Here is the uncomfortable pattern we keep seeing: a team ships an AI feature, usage spikes for two weeks, then flatlines, and six months later nobody can say whether it worked. The model was never the problem. The absence of a definition of success was.
Why teams start in the wrong place
Most AI initiatives begin with technology envy, not a problem: a competitor shipped something, a demo impressed a founder, and 'we should have AI too' becomes the entire business case. That starting point fixes the finish line at launch; almost nobody stops to ask what 'working' would even look like. Launches are easy to celebrate and impossible to learn from.
The three questions
One: whose time or money does this feature save, concretely? Two: which single metric will tell you it's working — and what number counts as failure? Three: what does the non-AI baseline achieve? If a rule-based shortcut or a better default gets you 80% of the way, ship that first and let it embarrass the model.
Write the metric first, touch the model second
We write the acceptance metric into the requirements before anyone touches a model, and we build the measurement plumbing before the feature. That order feels slow for exactly one sprint. Then it makes every decision afterwards — model choice, scope cuts, kill-or-scale — fast, because there is a number to argue with.
The verdict is simple: an AI feature that can't name its metric hasn't earned the right to be built.