AI tool retries: prevent a timeout from becoming a duplicate action
Use operation identities, durable state, reconciliation, and bounded retries when an agent calls tools that change the world.

A tool call times out. The assistant tries again. Two calendar events now exist, even though the user requested one. The first request may have succeeded on the remote service before its response was lost.
For agents that modify records or trigger workflows, a timeout is an ambiguous observation. It does not prove that nothing happened. Reliable recovery requires an operation model outside the language model's conversational memory.
Separate transport failure from business outcome
Imagine a fictional event-planning assistant that reserves a meeting room. The reservation service accepts the request, writes the booking, and then the network connection drops. The assistant receives no confirmation, but the room is already reserved.
Repeating the request with a new identity can create another booking. Asking the model to avoid duplicates is insufficient because it cannot observe the remote state reliably from the failed response alone.
AWS's Builders' Library discusses idempotent APIs as a way to make retries safer. The application-level design below applies that principle to agent tools; it is an illustrative architecture rather than a guarantee supplied by any particular model.
Give the operation a stable identity
Create a server-side operation identifier for the user's intended action before dispatch. Reuse it for retries of that same action. Where the downstream service supports idempotency keys, bind the key to the authenticated scope and the intended parameters.
Do not let the model invent a fresh key each time it becomes uncertain. A retry is the same operation; a changed room, time, or attendee set may be a different operation. Define that distinction in application code.
Reject reuse of an operation identity with conflicting parameters. Otherwise, a key intended to prevent duplicates can mask a different action. Keep retention and expiry rules aligned with the period during which delayed retries may still arrive.
Record state durably
The room-reservation workflow might track proposed, authorized, dispatched, confirmed, failed, and unresolved states. Store state changes in a durable system so a process restart does not erase the history.
Mark confirmation only when the service returns sufficient evidence or a reconciliation check establishes the outcome. The assistant can explain that a booking is being checked, but it should not convert uncertainty into a claim of success.
Prevent concurrent workers from dispatching the same operation independently. Use appropriate transactional or coordination mechanisms, and test the race rather than relying on one worker usually being faster.
Reconcile ambiguous outcomes
After a timeout, query the downstream service by operation identity or another reliable reference if supported. If the booking exists with matching parameters, return that result. If the service confirms absence, a bounded retry may be appropriate.
Some APIs provide no reliable lookup or idempotency support. In that case, the product may need human reconciliation or a deliberately restricted automation scope. Do not promise exactly-once behavior when the system cannot establish it.
Also distinguish compensating actions from retries. Canceling an accidental booking is a new operation with its own failure modes. It should not be hidden inside an explanation that implies the original attempt never occurred.
Bound retry behavior
Set a maximum attempt count, an overall deadline, and backoff appropriate to the service. Retry only failures classified as potentially recoverable. An authorization error or invalid request usually requires a correction, not repetition.
Coordinate retries across layers. If the HTTP library, tool wrapper, job queue, and agent each retry independently, a small limit at each layer can multiply into many remote calls. Assign one layer responsibility for the policy and make the others visible.
Cancellation needs an explicit meaning. A user closing the chat does not necessarily cancel a request already accepted by the booking service. Explain the operation's current state and provide a supported cancellation path where one exists.
Test the uncomfortable timing windows
Simulate a lost response after success, a worker restart before confirmation, two concurrent retries, and a delayed success arriving after a timeout. Verify both the number of bookings and the message shown to the user.
Keep a human-readable operation record with safe identifiers and outcomes. It should support investigation without storing unnecessary private content. An agent becomes dependable when its recovery behavior is designed around uncertainty, rather than assuming every missing response means another attempt is harmless.