Skip to content
All insights

AI decisions

When AI is actually useful, and when it isn't

A working test for whether a problem wants a model, a rule, a query, or a conversation with the person who owns the process.

· 6 min read · Deep Sharma

Most requests that arrive as “we need AI for this” are really one of four requests: we need to automate a judgement, we need to find something in a lot of text, we need to generate something that used to take a person an hour, or we need to look modern. Only the first three are engineering problems. The fourth is a marketing problem wearing an engineering hat, and it is worth naming out loud before any money moves.

The test we actually use

Before a model is on the table, we ask whether the problem has three properties. It needs to be genuinely fuzzy. If a human can write down the rule, write down the rule. It needs to tolerate being wrong sometimes, because a probabilistic system will be, and the cost of the wrong answer sets the design. And it needs enough of the right data to evaluate against. Not to train on; to evaluate on. Without an evaluation set you don't have an AI project, you have a demo.

  • Fuzzy: the input varies in ways a rule can't enumerate (free text, images, messy records).
  • Tolerant: a wrong answer is recoverable, reviewable or cheap; or the workflow keeps a human in the loop where it isn't.
  • Evaluable: you can assemble a few hundred real examples with known-good answers, and someone owns that set.

Where AI usually earns its place

Classification and routing of unstructured input. Extraction of fields from documents that were never designed to be parsed. Summarisation where the reader needs the gist and can open the source. Drafting where a person will edit before anything is sent. Search over content that keyword search handles badly. In each case the human stays in the loop where the cost of error is high, and the model does the part that used to be tedious.

Where it usually doesn't

Anything with a deterministic answer. Arithmetic, lookups, validation, business rules that are already written down. Anything where a wrong answer is expensive and unreviewable. Anything where the real blocker is that nobody owns the process. A model will not fix an organisational gap; it will automate the confusion. And anything where the operating cost of inference per transaction exceeds the value of the transaction, which is a calculation people skip surprisingly often.

If a human can write down the rule, write down the rule.

What this looks like on the evidence ladder

  • factSupport receives around N tickets a day; agents spend a measured amount of time triaging each (from the ticketing system).
  • assumptionMisrouted tickets are the main cause of slow resolution.
  • hypothesisA hosted model can route to the correct team at an accuracy that beats the current process on our data.
  • constraintTicket content includes customer data that must stay in-region.
  • decisionRun a two-week evaluation on 500 historical tickets before any integration work.
  • tradeoffTwo weeks spent measuring, not building, accepted because it decides whether the build is worth doing at all.

Notice that the decision isn't “use AI”. It's “find out whether AI would work here, cheaply, before deciding”. That is nearly always the right first decision, and it is the one that most projects skip.

Free · 30 minutes · one real problem

Bring a problem. Leave with clarity.

Thirty minutes, one real problem, structured thinking. If there's no value, there's no engagement.