Skip to content
All insights

AI operating cost

The AI cost that arrives every month

Inference, retrieval, evaluation, re-runs and the human review loop: how to model the operating cost of an AI feature before you build it.

· 6 min read · Deep Sharma

Traditional software has a build cost that dominates and a run cost that is mostly hosting. AI features invert this. The build can be quick. A prompt, a retrieval step and an evaluation harness get you a long way. Then a cost arrives with every single transaction, forever, and it scales with usage in a way most product economics were not designed for.

The components people forget

  • Input tokens dominate for retrieval-heavy features. Every document you stuff into context is paid for on every call.
  • Retries and fallbacks multiply cost. A validation failure that triggers a re-run doubles that transaction.
  • Evaluation is a recurring cost, not a one-off; every prompt or model change needs re-scoring against the reference set.
  • Human review is part of the run cost when the design keeps a person in the loop, and it should, where errors are expensive.
  • Model deprecation forces migration work on the vendor's schedule, not yours.

A model you can defend

Estimate transactions per month as a range, not a point. Measure the average tokens in and out from a realistic sample; don't guess them. Multiply through at current pricing and label it as an assumption, because pricing changes. Add the re-run rate observed in evaluation. Add the review cost where review exists. Then compare the per-transaction cost with the per-transaction value, and be honest when the second number is unknown.

  • factAverage context size per request, measured on a sample of real inputs.
  • assumptionUsage will be roughly proportional to current manual volume.
  • hypothesisA smaller model with retrieval matches the large model's accuracy on our evaluation set at a fraction of the cost.
  • constraintMonthly operating cost must stay within the agreed ceiling at projected volume.
  • tradeoffCaching and truncation reduce cost but can reduce quality on long documents, accepted with monitoring.

If the run cost per transaction exceeds the value per transaction, the feature does not become viable by being built well. That is a STOP, and it is a far cheaper STOP to reach on a spreadsheet than in production.

Free · 30 minutes · one real problem

Bring a problem. Leave with clarity.

Thirty minutes, one real problem, structured thinking. If there's no value, there's no engagement.