There is a well-worn pattern in AI projects. A demo goes well. Everyone is excited. Six months later the thing is either quietly switched off or held together by a person whose full-time job is now correcting its output.
The failure almost never happens at the model. It happens in the gap between "this works on the ten examples we tried" and "this works on the ten thousand we did not." These are the five questions we work through before writing anything, and what a bad answer to each one costs.
1. What does a wrong answer cost?
This is the first question because it determines everything else.
If a wrong answer costs a user three seconds of annoyance, you can ship at 85% accuracy and iterate. If a wrong answer means an incorrect invoice, a missed safety issue, or a customer told something legally binding and false, then 99% is not good enough either — you need a human in the loop, and the project is now about how to route the uncertain cases to a person rather than about the model at all.
Teams routinely skip this and end up building a fully automated system for a problem that needed an assisted one. The rework is expensive because the architecture is different: an assisted system needs a review queue, an audit trail, and a confidence signal the model was never designed to produce.
Bad answer cost: a rebuild, not a fix.
2. How will we know if it is working?
"It seems better" is not a measurement, and neither is a demo.
Before building, we want an evaluation set: a few hundred real examples with known-good answers, drawn from actual production data rather than invented. It is tedious to assemble. It is also the single highest-leverage thing on the project, because without it you cannot answer whether a prompt change helped, whether a model upgrade regressed something, or whether the system is degrading as your data drifts.
The teams that build an eval set first move faster within a month. The ones that skip it are still arguing about whether the last change helped in month four.
Bad answer cost: every subsequent decision becomes a matter of opinion.
3. Is the data actually available?
There is a specific and very common failure here. Someone says "we have all that data," and they do — spread across a CRM, a shared drive of PDFs, an inbox, and one person's head.
The question is not whether the information exists. It is whether it exists in a form a system can retrieve at request time, with correct access controls, at acceptable latency. Often the honest answer is that six months of data engineering sits between where you are and where the AI project can start. That is fine, but it should be a known cost, not a discovery in week five.
There is a security dimension too. If you are building retrieval over internal documents, the retrieval layer must respect the same permissions the source systems do. It is startlingly easy to build a search interface that cheerfully surfaces the salary spreadsheet to everyone.
Bad answer cost: the timeline doubles, and you find out late.
4. What is this costing per request, at volume?
Model pricing is easy to underestimate because the demo is cheap and the demo is not the thing.
Work out the real number: tokens per request including the system prompt and retrieved context, requests per user per day, users at the volume you are planning for. Then check whether that number is compatible with the margin on the product. We have seen features that were fine in testing and would have cost more per user than the subscription price at scale.
Usually there are answers — caching, routing simple requests to a smaller model, tightening what goes into context. But they are architectural decisions, and they are much cheaper made at the start.
Bad answer cost: a feature you have to withdraw after launch.
5. What happens when the model changes?
Models get deprecated. Providers change defaults. A prompt tuned against one version can behave differently against the next, sometimes subtly enough that nobody notices for weeks.
That means pinning versions explicitly, keeping the eval set runnable on demand, and putting a seam between your application logic and the specific provider so a migration is a contained piece of work rather than a rewrite. Not full provider abstraction — that usually costs more than it saves — but enough of a boundary that swapping is a known job.
Bad answer cost: an outage you did not cause and cannot immediately fix.
The one that matters most
If you only work through one of these, make it the second. An evaluation set turns every later argument into a measurement. Without it, you are shipping on vibes, and vibes do not survive contact with real users.
If you are weighing up an AI project and want an outside read on it, get in touch. We will tell you honestly if we think it is worth building.