AI agent architecture
AI agents: a case for restraint
Most “agent” problems are workflows with one or two model calls in them. Knowing the difference is the architecture decision.
An agent, in the useful sense, is a system where a model decides what to do next: which tool to call, whether to continue, when it is done. That autonomy is powerful when the path through a task genuinely varies. It is a liability when the path is known, because you have replaced a deterministic sequence with a probabilistic one and added cost, latency and failure modes to get there.
Workflow first
Start by writing the task down as steps. If the steps are the same every time, you have a workflow: fixed sequence, model calls where judgement is needed, ordinary code everywhere else. Most document processing, most extraction, most classification-then-route problems are workflows. They are cheaper, more testable and easier to explain to an auditor.
When autonomy is justified
- The set of possible actions is large and the right sequence depends on what earlier steps reveal.
- The cost of a wrong action is bounded: reversible, sandboxed, or reviewed before it takes effect.
- You can evaluate end-to-end outcomes, not just individual steps, on a representative set of tasks.
- There is a clear stopping condition and a budget (steps, tokens, time) that the system cannot exceed.
The NFRs nobody writes down
Agents need observability at the level of decisions rather than requests. What did it see, what did it choose, and why. They need idempotent tools, because retries happen. They need a cost ceiling per task. They need a human escalation path that is a first-class feature, not an error state. If those are not in the design, the design is not finished, regardless of how good the demo looked.
Replace a known sequence with a model's judgement only when the sequence is genuinely unknown.

