AI operating cost
The AI cost that arrives every month
Inference, retrieval, evaluation, re-runs and the human review loop: how to model the operating cost of an AI feature before you build it.
Traditional software has a build cost that dominates and a run cost that is mostly hosting. AI features invert this. The build can be quick. A prompt, a retrieval step and an evaluation harness get you a long way. Then a cost arrives with every single transaction, forever, and it scales with usage in a way most product economics were not designed for.
The components people forget
- Input tokens dominate for retrieval-heavy features. Every document you stuff into context is paid for on every call.
- Retries and fallbacks multiply cost. A validation failure that triggers a re-run doubles that transaction.
- Evaluation is a recurring cost, not a one-off; every prompt or model change needs re-scoring against the reference set.
- Human review is part of the run cost when the design keeps a person in the loop, and it should, where errors are expensive.
- Model deprecation forces migration work on the vendor's schedule, not yours.
A model you can defend
Estimate transactions per month as a range, not a point. Measure the average tokens in and out from a realistic sample; don't guess them. Multiply through at current pricing and label it as an assumption, because pricing changes. Add the re-run rate observed in evaluation. Add the review cost where review exists. Then compare the per-transaction cost with the per-transaction value, and be honest when the second number is unknown.
- Average context size per request, measured on a sample of real inputs.
- Usage will be roughly proportional to current manual volume.
- A smaller model with retrieval matches the large model's accuracy on our evaluation set at a fraction of the cost.
- Monthly operating cost must stay within the agreed ceiling at projected volume.
- Caching and truncation reduce cost but can reduce quality on long documents, accepted with monitoring.
If the run cost per transaction exceeds the value per transaction, the feature does not become viable by being built well. That is a STOP, and it is a far cheaper STOP to reach on a spreadsheet than in production.

