When a tool call fails, an AI agent should first work out what kind of failure it is. Temporary faults such as timeouts or rate limits can be retried a few times, but only if the action is safe to repeat. Errors caused by bad input should be fixed, not retried. Anything else should stop and go to a person.
Why tool failures matter more for agents
Every integration fails sometimes. APIs time out, services go down for maintenance, tokens expire and rate limits kick in during busy periods. Conventional software handles this with code a developer wrote for each case. An agent is different: it decides for itself what to do next, and without clear rules it tends to do one of two unhelpful things.
The first is to retry blindly. If a payment, order or email tool timed out after the action actually went through, a retry can do it again. The second is to improvise: the agent reads an error message, guesses at a fix, and tries a different tool or different arguments that nobody tested. Both can turn a small, recoverable fault into a real incident. Analysis of agent failure modes, such as Openlayer's review of tool-calling errors and loops, shows how one bad call can spread through the rest of a task.
Sort failures into kinds
The agent cannot choose a sensible response until it knows what went wrong. Most tool failures fall into a small number of groups, and the tool layer should report which group applies rather than leaving the model to interpret raw error text.
Temporary faults
Timeouts, dropped connections, rate limits and brief service outages usually clear on their own. These are the only failures where a retry is the right first move, and even then with a short wait that grows between attempts.
Input and permission errors
A rejected request because a field is missing, a value is out of range or the agent lacks permission will fail the same way every time. The agent should correct the input if the error says exactly what is wrong, or stop and report it. Retrying the identical call only adds noise, and a permission error should never prompt the agent to look for another route to the same data, as covered in stopping agents reaching restricted data.
Unknown outcomes
The hardest case is when the agent does not know whether the action happened, for example a timeout on a call that creates a record. Here the agent should check the state of the target system before doing anything else, rather than assume the action failed.
Make actions safe to repeat
Retries are only safe when repeating a call cannot cause a second effect. For tools that read data this is usually true already. For tools that create, change or send something, it has to be designed in.
The common pattern is an idempotency key: a unique reference generated for each intended action and sent with the request. The receiving system records it and, if the same key arrives again, returns the original result instead of acting twice. Many payment and messaging APIs already support this, and internal APIs can add it at the integration layer without changing the systems behind them. Where a tool cannot support it, the agent should look up whether the action already happened before trying again.
This is integration work more than model work. The agent's prompt can say "do not repeat actions", but only the tool design can guarantee it.
Put limits around retries
Even safe retries need a ceiling. Set a maximum number of attempts per call, a maximum number of failed calls per task and an overall time budget. When a tool keeps failing across many tasks, a circuit breaker should stop calls to it for a while so the agent is not hammering a service that is already struggling. These limits also keep costs predictable, which ties into controlling what an agent costs to run.
Fallbacks are worth defining in advance: a read-only alternative source, a cached answer that is clearly marked as such, or a simpler path that does part of the job. What matters is that each fallback is chosen and tested by the team, not invented by the model at run time.
Know when to stop and hand over
Some failures should end the attempt. When retries are exhausted, when the outcome of an action is still unknown, or when continuing would mean acting on stale or incomplete information, the agent should stop, leave the task in a clear state and pass it to a person with a short account of what it tried. A well-designed human in the loop step makes that handover quick rather than disruptive.
Every failure, retry and handover should also be logged with the tool, the arguments, the error and what the agent did next. That record is what lets you audit what an agent did, spot tools that fail often and monitor the agent in production with real signals rather than guesses.
Test failure paths before launch
Most agent testing checks what happens when tools work. Failure handling needs its own tests: simulate timeouts, rate limits, malformed responses and partial successes, then check that the agent retries only what it should, never duplicates an action and stops cleanly when it must. These cases belong in the same evaluation suite described in what AI agent evals catch, and they are among the things that break first in production when they are skipped.
CodeDTX's enterprise AI integration and AI agent development teams design tool layers that classify failures, make actions safe to repeat and hand over cleanly. If your agents are already calling production systems, talk to CodeDTX about reviewing how they handle failure.
Frequently asked questions
Should the model or the code decide whether to retry?
The code should decide in most cases. The tool layer knows whether an error is temporary, whether the action is safe to repeat and how many attempts remain, so it can apply those rules consistently. The model is better placed to decide what to do once the rules say a retry is not allowed, such as asking for missing information, trying an approved fallback or handing the task to a person.
What is an idempotency key in plain terms?
It is a unique reference attached to a request that says "this is one specific action". If the same request arrives twice with the same reference, the receiving system recognises it and returns the first result instead of acting again. For an agent, that means a retry after a timeout cannot create a second order, payment or message, even if the first attempt actually succeeded.
How many times should an agent retry a failed tool call?
There is no single right number, and it should be set per tool rather than left to the model. Read-only calls to a service with occasional timeouts can tolerate a few attempts with growing waits between them. Calls that change data should retry only when they are safe to repeat, and a failure that keeps recurring should trip a circuit breaker and go to a person.
What should the agent tell the user when a tool fails?
It should say plainly what it was trying to do, that the step did not complete and what happens next, whether that is a later retry, an alternative route or a handover to a colleague. It should not expose raw error messages or internal system details, and it should never claim an action succeeded when the outcome is unknown. Honest, short updates keep trust even when things go wrong.



