Short answer
#Three in the morning
The most expensive automation looks boring. It does not crash and it raises no errors. It carefully carries one wrong assumption all the way to the end.
The first step reads incoming mail and marks one message as a request to close an account. The confidence behind that judgement is low, but what travels to the next step is an ordinary string: request type, closure. The second step cancels the renewal. The third queues a data-deletion job. The fourth emails the customer to confirm the closure. The fifth pushes the status into the accounting system.
By 03:10 all of it is done. The customer had written about changing plan. Six systems now agree that they have left, and each of them is confident, because it received the fact from the one before.
There was one mistake and it was small. What made it an incident was not the mistake but the four trusting steps after it. Note that each of those steps is written correctly and does exactly what it is supposed to do.
#Why a chain amplifies a mistake
The mechanism is simple. The first step knew it was unsure. The second step did not: what reached it was a result, not an assessment of a result. From there confidence only grows, because every further step sees consistent data from several sources, and consistency looks like corroboration.
The cheapest places to stop are two: before the step that leaves the company, and before the step that cannot be taken back.
One practical conclusion follows. Risk has to be assessed across the whole chain rather than per tool call. The question of what this step does is a safe question. The question worth asking is different: what happens to the entire chain if the first step was wrong, and at which point can that still be stopped.
#Where a chain breaks
Stopping points do not go after every step, or the automation loses its purpose. They go on the crossings between classes of action, and there are only three.
- From reading to writing. While the chain reads and computes, a mistake costs electricity. The first write makes it visible to others.
- From internal to external. An email, a message, a webhook, a record in somebody else system. There is no undo here, only an apology.
- From reversible to irreversible. Money, deletion, permission changes, publication. You always stop here, even when stopping is slow.
A useful exercise: take your longest automation and mark those three crossings on it. It usually turns out they happen earlier than expected, and that there is not a single check between them.
#Depth, retries and budgets
Three numbers that have to be stated explicitly. Not because they prevent a mistake, but because without them the chain has no end.
-
Maximum depth
How many steps one run may take. A chain with no limit eventually loops: the agent sees the result of its own previous action, treats it as new input and carries on.
Hitting the limit is a signal, not an error. It has to lead to a stop and a record rather than to a silent retry.
-
A retry budget
How many times a step may be repeated and at what interval. Unlimited retries are not resilience; they turn one unavailable external service into a thousand requests and a set of duplicated operations.
When the budget is spent, the work goes to a person for review rather than dissolving.
-
A tool-call budget
How many actions the agent may take in one run in total, and how many of them may change something. It is the same volume limit as the cap on bulk operations, applied to length instead of width.
Log the remaining budget as well as the calls: that graph shows chains getting longer long before an incident does.
#The same step twice
Any chain with retries will eventually run one step twice. The network dropped after the request was sent and before the reply arrived. The job was restarted. A person pressed the button again. The question is not whether this happens but what comes of it.
Idempotency means a repeated step produces no second consequence. In practice it is usually an operation key: the caller invents a unique identifier, sends it with the request, and the receiver promises that a second request with the same key creates nothing new and returns the result of the first. Payment providers have supported this for years, for exactly this reason.
A one-minute test
Take the step that sends an email or moves money and run it twice with identical input. If you get two emails or two charges, there is no idempotency, and the retries in your chain are working against you.
#A re-check before the irreversible
The cheapest item on the list and the rarest. Before an irreversible step, the critical fact is established again, and by a different route than the one that produced it.
In the scene at the top of this page one check would have been enough: before sending the closure email, look for a confirmed customer action to close, rather than for a classifier output. The check costs one query, and it breaks the chain at the point where the damage has not yet left the company.
By a different route is not decoration. Re-checking with the same classifier and the same prompt returns the same answer and manufactures a feeling of corroboration. What works is a deterministic check: a field in the database, an order state, a signature, a user action, a reply from an external system.
#Step type and how to stop it
| Type of step | What stops it | What undoes it |
|---|---|---|
| Classification or extraction | A confidence threshold; low confidence routes to a person instead of onward | Nothing to undo, provided the next step has not begun |
| A write to an internal system | An operation key, a short transaction, a cap on the number of objects | A reverse write recording the reason and the previous value |
| Queueing a job | A delayed start with a cancellation window; the job re-reads state before acting | Cancellation, while the job has not started |
| A message leaving the company | Re-checking the fact by another route, a delay of some minutes, confirmation | Nothing: only a second message with a correction |
| Money | Confirmation by a person, a cap on the amount per day | A compensating operation and a record in the log |
| A write into somebody else system | A circuit breaker on the number of anomalies inside a time window | Only a compensating operation, and only if their API has one |
The circuit breaker in the last row is worth building separately from everything else. The rule can be crude: if the chain made more than N changes in ten minutes, or more than N per cent of them are of one kind, it stops and waits for a person. A crude rule that works at night is more useful than a precise one somebody promised to write.
#Reviewing a long chain
Six questions to ask before an automation runs
Walk your longest chain step by step with a pen.
- Where does uncertainty enter this chain, and does it travel onward in any form at all? If the next step receives only a result, the whole chain treats it as a fact.
- How many steps and retries are allowed, and what happens when the limit is reached? A limit with no handling is a quiet stop nobody hears about.
- Which steps survive being run twice without a second consequence? That is the idempotency test, and it is usually failed.
- Which fact is re-checked before the first irreversible step, and by what route? A re-check by the same route does not count.
- What can be undone afterwards, and who will do it at three in the morning? A compensating operation has to be written in advance, not invented on the night.
- What rule stops the chain without a person, and on what signal? At night the decision is made by a threshold somebody thought to set, not by the team.
#How we do it here
1ADK is itself a chain: the agent completes units of work, submits results, and the platform verifies and stores them. So the same rules apply inside it. Processing the same package twice is safe and does not create a second copy of the knowledge. Unfinished files live apart and are only activated after they have been validated, and the last known-good generation is never destroyed first.
And one more thing that belongs to this subject directly. What is unknown stays unknown here: if the agent could not establish something, that is recorded as a gap rather than promoted to a confident statement. The coverage of an analysis decides what may even be reported as gone. It is the same idea as this article, applied to knowledge: uncertainty has to reach the next step together with the data.
#What this does not give you
What this does not do
- Transactions are not available everywhere. In external services a compensating operation is often the only option, and sometimes there is not even that.
- A re-check costs time and money, so it goes on the irreversible steps rather than on all of them.
- Stopping points reduce autonomy. That is a real trade rather than a free improvement.
- None of these limits fixes a wrong classification. They bound what grows out of one.
- A circuit breaker nobody tuned against real traffic will fire on a peak day and teach the team to switch it off.
Take your longest automation and count the irreversible actions it can complete while nobody is watching. Which of them will you leave automatic once the cost of being wrong is written next to it?