SaaSexample use case
From AI prototype to production — rescuing a stalled project
A contractor built a working demo and left. The internal team tried to develop it, but every prompt change is a lottery and the API bill keeps growing. We do not start by rewriting — we start by making it measurable what is better.

Sound familiar?
- The prototype works “mostly”, but nobody can say on what share of cases, or where it fails.
- A prompt or model change fixes one case and breaks three others — and you hear about it from users.
- Token costs are growing and nobody knows which feature is consuming them.
- The model provider announced a version deprecation and you have no way to verify the new one behaves the same.
How it works

1. Audit what you have
Code, prompts, data, integrations, costs. The output is a list of risks and an estimate of what can be saved and what is cheaper to rebuild.
2. A test set from real cases
We collect 100–300 real inputs with the expected output. From now on every change has a number — better or worse, not “it seems”.
3. Monitoring and cost limits
Logging of every call, cost per feature and per user, alerts on deviation. The budget gets a ceiling that cannot be exceeded silently.
4. Optimisation and handover
Only now do we change prompts, the model or the architecture — with measurement. At the end we hand the project to your team with documentation, or we run it.
What the system handles
- Prototypes on OpenAI, Anthropic, Google, Mistral and open-source models; LangChain, LlamaIndex, custom code.
- Chatbots, RAG systems, data extraction, classification, agents with tools.
- Prompt and model versioning with rollback.
- Migration to a cheaper or smaller model where the test set shows it does no harm.
- EU AI Act readiness: documentation, logging, transparency towards users.
Human oversight and safety
- No change goes to production without running the test set.
- Gradual rollout: a new version first for a share of users, with metric comparison.
- Hard cost limits per day and per user; alert on anomaly.
- Code, prompts, test set and data stay yours — with no dependency on us.
What result is realistic
evaluations
instead of guessing whether it improved
After 4–6 weeks you have numbers instead of a feeling: success rate on the test set, cost per case and a list of where the system fails. Cutting token costs by tens of percent is common, simply by using a cheaper model where measurement shows it is enough.
The figures are indicative for typical volumes. We give a precise estimate for your company after analysing the process.
Indicative scope
- Audit of the existing solution
- €1,500 – €3,000
- Test set, monitoring and cost limits
- €3,500 – €7,000
- Optimisation and completion into production
- from €5,000
- Operations and monitoring
- from €390 / month
Prices are indicative and exclude VAT. We give an exact figure only after analysing your process — not off the cuff on the first call.
Frequently asked questions
Is it not cheaper to start over?
Sometimes yes, and the audit says so within two weeks. Surprisingly often, though, more can be saved than expected — the problem is usually missing measurement, not the code.
Do you work with code your team did not write?
Yes, that is the essence of this service. We need access to the code, the prompts and a sample of real inputs. An NDA goes without saying.
What is a “test set” and why does it come first?
A list of real inputs with what the system should return. Without it, nobody can say whether a change helped. With it, every tweak is a decision on numbers — and that is the difference between a demo and production.
Will you take over operations too?
Your choice. We can hand the system to your team with training, or run it with monitoring and a monthly report.
Related service
From prototype to production
We build this solution as part of our service „From prototype to reliable production“.
Got a working demo you cannot get into production? This is exactly where most AI projects die.Read more
- From PoC to production: why most AI projects never shipAI projects do not die because the model fails. They die in the space between “it looked good in the demo” and “it runs reliably and we know what it costs”.
- MCP: the protocol that finally made AI a usable toolThe Model Context Protocol is roughly to AI what USB was to peripherals. An explanation without the marketing — and what follows from it when you are considering an integration.
More solutions
WholesaleAutomating email orders into your ERP with AI
~70% · of orders handled without a human
ManufacturingCompany document search — an answer with its source in seconds
2–4 h · saved per person per week
AccountingAutomated invoice and delivery note processing with AI
< 3 months · typical payback periodIf you have a demo that is afraid of production, the first step is cheap: an audit that says what can be saved.
You do not need to know whether you need AI, automation or a new system. Show us the process that slows you down — we will tell you what can be automated and whether it pays off.