Idempotent Tool Design for Retryable Agent Actions
Safely retrying agent actions requires idempotent tool design across three layers.

Idempotent tool design is the discipline that decides whether an AI agent can safely retry a failed action. A hotel booking agent times out mid-transaction, retries the charge, and the guest's card gets billed twice because nothing in the stack recorded that the first charge already went through.
Why agent retries create duplicate side effects
The core failure sits in the retry logic itself: agent frameworks retry at the level of the language model's output, and nothing at that level tracks whether the tool call already fired and changed something in the world. LangChain, LlamaIndex, OpenAI's Agents SDK, and Anthropic's SDK all build in automatic retries for transient failures, and that's the right move for any distributed system. A timeout means the caller didn't hear back, not that nothing happened.
The tool, meanwhile, may have already written the database row, charged the card, sent the email, or opened the support ticket. The framework sees silence and assumes failure, so it tries again. But the delivery systems underneath these agents are built to deliver at least once, not exactly once: Amazon's SQS standard queues state plainly that duplicate delivery is part of the contract, and HTTP retries work the same way by design. The agent stack inherited this behavior without inheriting the decades of engineering built around it to keep duplicates from turning into damage.
The same pattern occurs across industries, not just one. In each case, a timeout or a transient error triggers the failure, not a flaw in the agent's reasoning. The gap is structural: nothing downstream of the model call knows what already happened, so the agent has no way to find out before it acts again. The LLM compounds this structural gap: it produces roughly the right tool call most of the time, but a small fraction of the time it produces an almost-but-not-quite duplicate, and at scale that fraction is enough to expose every non-idempotent operation in the downstream stack.
How LLM non-determinism turns a distributed-systems problem into a vulnerability
Every standard fix for this kind of problem, unique request IDs, recorded values from past runs, deterministic logic in the orchestrator, rests on one assumption: that a retried request looks exactly like the original request. LLM agents break that assumption by design.
Even at temperature zero, floating-point rounding inside GPU kernels can produce a different sequence of tokens on a second run, so an agent restored from a checkpoint can generate a request that looks brand new to the server receiving it. A paper called ACRFence, published in 2026, gives this failure a name: the semantic rollback attack. The paper sorts this into two attack classes: Action Replay, where an irreversible effect happens twice, and Authority Resurrection, where a credential that was already spent gets reused after the agent is restored.
This shows up across the agent ecosystem, not in one framework. A review of 12 major agent frameworks found that none of them enforce exactly-once behavior where the agent calls a tool. In LangGraph, tools can fire again on resume, and a maintainer confirmed the issue directly, writing that fixing it would require tracking interrupt IDs, something the current stored metadata can't support. Google's ADK documentation states outright that rewinding an agent cannot undo whatever already happened outside the system. A GitHub issue on Claude Code, numbered 13897, documents an agent calling the same tool twice after a user approves a permission request, running the action redundantly. OpenClaw had a webhook replay bug that let the same event get processed more than once, serious enough to warrant a GHSA security advisory.
Checkpoint-restore gets sold as a reliability feature, a way to pick an agent back up after a crash without losing its place. Fixing that can't happen inside the framework's retry loop, because the retry loop is the thing generating the new request in the first place.
MCP failure results and the recovery information agents lack
When a client gets back a result marked isError:true, it knows something went wrong. It often has no further basis, nothing machine-readable, to decide whether to fix an argument, re-authenticate, wait and retry, switch to a different tool, or stop.
MCP is better at describing that a failure happened than at describing what to do about it. Its typed fields can expose that something failed, and sometimes offer a broad policy hint, but they don't expose a specific cause, a specific target to repair, a concrete executable fix, or any constraint on whether replay is safe. The plain-text description inside an error result often carries more useful detail about cause and target than the typed fields do, but that's exactly the problem: it means a client has to interpret natural language to recover safely, which is precisely the kind of judgment call deterministic retry logic isn't built to make.
A retry is never a neutral move. Retrying a malformed request just burns a call. Retrying an operation that partially committed can cause real damage. Other protocols treat this as a solved design question. An open issue on the MCP spec now proposes a structured, schema-governed error object to close exactly this gap, which signals that the protocol's maintainers see the problem but haven't resolved it yet.
Because the protocol layer can't tell an agent what to do after a failure, that information has to live somewhere else: in the contract of the tool itself.
Idempotency requirements across the tool call stack
Making a non-idempotent action safe to retry takes deliberate design at three layers that all have to cooperate: the agent runtime, the tool execution layer, and the tool interface.
The agent runtime has to generate an idempotency key for each step in a workflow and hold onto it. The key needs to come from something durable, a pattern like {workflowRunId}:{stepId} works well in practice, so that the same key gets regenerated when the agent resumes rather than a fresh one getting minted every time it retries.
If the key already exists and the earlier call succeeded, the layer returns the cached response and skips execution entirely.
The tool interface itself has to accept the key and actually enforce deduplication on it. Tools that call external APIs need to pass the key through to that API. Tools that write to an internal database need to use the key as part of a unique constraint on the write.
A request ID is not the same thing as an effect ID. Idempotency, in other words, isn't a single on/off switch. It's scoped, time-limited, and specific to the operation it protects.
The ghost-order pattern names what happens when only one of the three layers holds. An order-capture agent called a create_order tool on an order management system, and that tool really was idempotent, as long as the client supplied the same order ID on every call. The tool's contract was correct. The agent's behavior defeated it anyway, because correctness at one layer means nothing without key reuse enforced at the runtime layer and prompt design that reinforces it.
Classifying tool operations to match retry policy to risk
Not every tool call carries the same risk when it's retried, and applying one blanket policy, always retry or never retry, gets the wrong answer in both directions. A read-only lookup can be retried as often as needed, since reading a record never changes it. A genuinely idempotent write, one that produces the same result no matter how many times it's called with the same input, can be retried freely within its documented contract, no extra machinery required. A high-consequence, non-idempotent action, a payment, an irreversible state change, needs a durable record of whether it already ran, and an ambiguous failure on one of these should trigger a reconciliation step rather than leave the model to guess what happened.
A design proposal under active discussion in a GitHub issue formalizes this into four classes: SAFE_RETRY, DEDUPLICATE, REQUIRES_EFFECT_CHECK, and NON_RETRYABLE, where each operation declares which class it belongs to, along with a deduplication scope, a retention window, and a way to query its status. Its candidate outcome states go beyond success and failure: ACCEPTED, RUNNING, COMPLETED, REJECTED_DUPLICATE, NOT_FOUND, FAILED_BEFORE_INPUT, and EFFECT_UNKNOWN. A NON_RETRYABLE operation that ends in doubt returns EFFECT_UNKNOWN rather than a false, falsely reassuring "it didn't happen." An unknown outcome can't be relabeled as FAILED_BEFORE_INPUT unless there's actual evidence to back that up, and a deduplication window that's expired can't quietly let an old command through as if it were a brand-new, valid one.
The real shift this framework proposes is the outcome-query contract: instead of retrying blindly, the planner asks the system what already happened and gets a status back. That replaces "I think I already did this" reasoning with an actual, checkable record. A second attempt carrying the same intent_id doesn't get to replay a DEDUPLICATE operation automatically; a new attempt only becomes valid after a status check comes back, or after a fresh check confirms what effect, if any, actually occurred.
None of this works, though, until the operation's class is known in advance. Classification determines retry policy.
A planning/execution split as the structural home for idempotency enforcement
The Planner takes the high-level goal and produces a deterministic, structured list of steps, producing no side effects. The Executor carries out the plan, and if a step fails partway through, that failure gets reported back to the Planner for a new plan, without corrupting whatever execution state already exists.
The payoff appears on retry. The LLM never gets called again for a step that already ran. The cached decision from the first pass gets replayed instead, and the execution layer, which already knows how to check for existing effects, handles the actual side effects safely. Because the model is never re-asked, its non-determinism simply doesn't enter the picture on a retry.
Apache Airflow built this into a flag: setting durable=True on its AgentOperator means cached steps replay on retry instead of running again, so an agent that made several LLM calls and then failed on the last one produces zero repeated LLM calls and zero repeated tool calls when it picks back up. Production systems built on patterns like this tend to combine all three protections at once: idempotency keys guarding individual tool calls, deduplication tables guarding internal actions, and checkpointing guarding the workflow as a whole.
The ACRFence mitigation from earlier also fits here, working across frameworks rather than inside any one of them. It keeps a record of irreversible tool effects as they happen, and when a checkpoint-restore occurs, it enforces a strict replay-or-fork rule: either the system replays the exact recorded effect, or it forks into a new, clearly distinct path, rather than letting the model re-synthesize a slightly different request and pass it off as the same one. That closes the loop on the semantic rollback attack directly: the agent no longer gets the chance to generate a new UUID for an action that was already recorded as done.
Saga patterns and compensating actions for multi-step workflows
A single idempotency key solves the problem for one tool call. A multi-step workflow, book the flight, reserve the hotel, charge the card, needs something broader, because a failure three steps in doesn't just need a safe retry. It needs a way to undo the steps that already succeeded.
This is the saga pattern: a workflow is modeled as a sequence of local steps, and every step that has a side effect gets paired with a compensating action that can undo it if a later step fails. If the workflow fails at the charge step, the saga doesn't leave the flight and hotel bookings stranded. It runs the compensating actions for each completed step, in reverse order, until the system is back to a consistent state.
None of this works without the idempotency groundwork laid out earlier. A compensating action is itself a tool call, and it inherits every risk a forward action does: it can get retried, it can fire twice, it can run into the same ambiguous-failure problem described in the MCP section. The classification framework applies here too. A compensating action that cancels a reservation is a different risk class than one that reverses a wire transfer, and the saga's orchestrator needs to know which class it's dealing with before deciding how aggressively to retry a compensation that itself fails.
The cleanest structural solution is separating the agent into a Planner that produces a deterministic, structured step list and calls no tools, and an Executor that works through the list sequentially and owns all idempotency enforcement. The Executor, which already owns idempotency enforcement for forward actions, extends naturally to own the saga's rollback path too: when a step fails, the Executor doesn't ask the model to figure out what to undo and how. It walks the list of completed steps in reverse and runs the compensating action already attached to each one, using the same idempotency keys and status checks that protected the forward path.
The result is a workflow that can fail safely at any point, beyond just a tool call that can retry safely at any point. A durable record of what happened, who committed what, and what hasn't been undone yet is what turns a chain of individually idempotent actions into a workflow that's trustworthy as a whole.
Sources
- The Idempotency Problem in Agentic Tool Calling - TianPan.co
- ACRFence: Preventing Semantic Rollback Attacks in Agent Checkpoint-Restore
- Define idempotent and uncertainty-safe retry semantics for agent actions · Issue #24 · Unjuno/agent-interface
- MCP: agree an interoperable wire contract for tool refusals, tool failures and overload (JSON-RPC codes, `data.reason`, `isError`, retry/replay) · Issue #2250 · terrene-foundation/kailash-py


