InsightsAI & Automation
What actually breaks when you put an AI agent in production
An agent that works in a demo and an agent that survives production differ in about six places. None of them are the prompt.
- Published
- Reading time
- 9 min read
- Topic
- AI & Automation
- Written by
- The Apoliums Team

An agent that works in a demo and an agent that survives production differ in about six places, and none of them are the prompt. The demo fails because the model said something wrong. Production fails because a tool returned a 200 with an empty body, the agent believed it, and the loop ran nineteen more times before the token budget stopped it.
Apoliums has shipped agents into scheduling, document review and internal support workflows. The list below is what actually broke, in the order it usually breaks.
What is the first thing that fails in a production agent?
Tool calls, not model output. An agent's tools are ordinary HTTP calls to systems that time out, rate-limit, return partial data, and occasionally succeed with a body the agent cannot use. The model has no way to distinguish "the API said there are no results" from "the API failed and returned an empty list", so it reasons confidently over an absence.
Every tool an agent can call needs three things the model never supplies:
- A timeout. Not the default. A specific one, shorter than the agent's own budget, so a hung dependency cannot consume the whole run.
- A typed result that distinguishes empty from failed. Return a discriminated result, never a bare array.
- An error the model can act on.
"upstream timeout, safe to retry once"changes agent behaviour."Error: 500"does not.
type ToolResult<T> =
| { status: "ok"; data: T }
| { status: "empty"; reason: string }
| { status: "failed"; retryable: boolean; message: string };
async function findBookings(patientId: string): Promise<ToolResult<Booking[]>> {
const res = await fetch(url, { signal: AbortSignal.timeout(4000) }).catch(
() => null,
);
if (!res) {
return { status: "failed", retryable: true, message: "upstream timeout" };
}
if (!res.ok) {
return {
status: "failed",
retryable: res.status >= 500,
message: `upstream ${res.status}`,
};
}
const data = await res.json();
return data.length
? { status: "ok", data }
: { status: "empty", reason: "no bookings for this patient" };
}
That shape costs twenty lines and removes the single largest category of agent misbehaviour, which is an agent hallucinating a plausible answer over a failed read.
Why do agents loop, and what stops it?
Agents loop because the stopping condition is a judgement the model makes about its own work, and a model that cannot complete a task will often decide to try the same approach again. Prompting it not to loop does not work reliably. Counting does.
Three limits belong in the runtime, not in the prompt:
- Maximum steps per run. A hard integer. When it trips, the run ends with a defined failure state that a human or a fallback path can pick up.
- Maximum repeats of the same tool call. Hash the tool name plus its normalised arguments. Two identical calls is a retry; four is a loop.
- A wall-clock budget. Steps can be cheap and still take minutes if each one waits on a slow dependency.
The important design decision is what happens at the limit. An agent that stops silently is worse than one that loops, because a loop is visible in the bill and a silent stop is only visible in the support queue. Emit a specific terminal state — step_limit_exceeded, repeated_tool_call, budget_exhausted — and route each one somewhere.
How do you test something non-deterministic?
You test the parts that are deterministic, and you evaluate the part that is not. These are different activities and conflating them is why teams end up with no confidence in either.
The deterministic parts are larger than they look. Tool implementations, argument validation, the retry policy, the loop guards, the parsing of structured output, the redaction of sensitive fields, the cost accounting — all of it is ordinary code and takes ordinary unit tests. If an agent's behaviour surprises you in production, the cause is in this layer far more often than in the model.
The non-deterministic part needs an eval set: a fixed collection of inputs with expected properties, run on every prompt or model change. Properties, not exact strings. "The response contains a booking reference", "the response never invents a doctor name not present in the retrieved context", "the tool sequence includes a lookup before a write". Score the set, store the score, and refuse to ship a prompt change that moves it down. Twenty cases you actually maintain beat two hundred that nobody reruns.
Why does the bill scale with retries rather than usage?
Because a retried run is a full run. When an agent fails at step seven of nine and the whole task restarts, the first six steps are paid for twice, and they were the expensive ones — long context, large retrieved documents, full history replayed at every step.
Three things control this, in order of effect:
Cache the deterministic prefix. System instructions, tool definitions and stable context are identical on every step of every run. Providers that support prompt caching charge substantially less for a cache hit than for fresh input tokens; whether you use that or not, the prefix must be byte-identical to be cacheable, which means no timestamps, no request IDs, and no shuffled tool ordering.
Do not replay the full history on every step. An agent at step twelve does not need the raw output of step three. Summarise completed sub-tasks into a compact result and drop the intermediate reasoning. Context grows linearly with steps unless something deliberately truncates it, and cost grows with context on every subsequent call.
Make retries resumable. If step seven fails, retry step seven. That requires the run to have durable state — completed steps, their outputs, the current plan — persisted outside the process. It is the difference between a retry costing one step and costing nine.
What does idempotency mean for an agent?
It means an agent that calls the same write tool twice must not produce two effects. This matters more for agents than for ordinary services, because retrying is the agent's default response to ambiguity, and ambiguity is its normal state.
Every tool with a side effect takes an idempotency key derived from the run and the intent, not from a random value — the same contract payment APIs have used for years:
const key = `${runId}:${stepId}:refund:${invoiceId}`;
await issueRefund({ invoiceId, amountMinor, idempotencyKey: key });
The server stores the key with the result. A second call with the same key returns the first result instead of performing the action again. Without this, a network timeout on a successful write — the single most common distributed-systems failure — turns into a double refund, a double email, or a double booking.
What should an agent never decide on its own?
Anything irreversible, anything financial, and anything that leaves the building. Apoliums treats these as a fixed category rather than a per-project judgement:
- Sending an external message to a customer
- Moving money, issuing credit, or changing a price
- Deleting data or revoking access
- Writing to a system of record that has no undo
For each of these the agent proposes and a person approves. The approval step is also the fastest source of training signal you will get, because the edits a reviewer makes are a labelled dataset of exactly where the agent is wrong. Teams that skip the review step usually skip it before they have any evidence about the failure rate, and then have no way to measure one.
What should you build first?
Build the observability before the second capability. Specifically: log every step with its inputs, outputs, tool results, token counts and latency, keyed by a run ID that appears in the user-facing interface. When somebody reports that the agent did something strange, the entire investigation should be one lookup.
The teams that get agents working are not the ones with better prompts. They are the ones who can answer "what did it actually do at 14:22 yesterday" in under a minute. Everything else in this article is downstream of that.



