Skip to main content

LLM provider outages: how to keep an AI agent working

When the model provider behind an AI agent slows down or goes offline, every task that depends on it stalls. Here is how to plan fallback models, test that they produce output your systems accept, and decide what the agent should do when no model is available.

Share -
An engineer in a white shirt sits in a dimly lit server room, hand to chin, watching monitors that show a glowing blue neural network brain and lines of system code.

To keep an AI agent working through an LLM provider outage, route model calls through one layer that can switch to a tested fallback model, detect failure quickly with timeouts and circuit breakers, and validate every fallback response against the same output checks as the primary. Then decide in advance what the agent does when no model can answer.

Why an outage hits agents harder than chat

A chat assistant that fails shows the user an error, and the user tries again later. An agent is usually partway through a task when the model stops responding. It may have read records, called tools and planned its next steps, and a single failed call can leave that work half done. Agents also make many model calls for one piece of work, so a short period of errors or slow responses touches far more tasks than the same period would in a simple chat product.

Outages are rarely clean, either. A provider may return rate-limit errors for one model, slow responses in one region, or partial failures that come and go. Your design has to handle the degraded cases, not only the full stop.

Put every model call behind one routing layer

If model calls are scattered through the codebase, switching providers during an incident means a code change under pressure. Send every call through a single gateway or client library that owns the provider list, credentials, timeouts and retry rules. That layer becomes the one place to change the model, and the one place to log which model answered each request.

GitLab publishes a runbook for failing over its AI gateway to another LLM provider, based on a runtime model-selection setting rather than a redeploy. Whether the switch is automatic or operator-driven, the principle is the same: changing the model should be a setting, not a release.

Detect failure quickly

Set timeouts per call type

A short classification call and a long planning call need different timeouts. Without them, the agent waits on a stalled request while the task it serves goes stale.

Use retries with care

Retry transient errors with backoff and a strict limit. Unlimited retries during an outage add load to a provider that is already struggling and hold your own queues open.

Add a circuit breaker

When errors from one provider pass a threshold, stop sending it traffic for a period and route to the fallback. A periodic trial request tells you when the primary has recovered, so traffic can move back in a controlled way.

A fallback model must pass the same checks

The most common failure in a failover is silent. The backup model answers, but it formats structured output differently, follows tool schemas less closely or handles long context in another way. Nothing errors, yet downstream systems receive output they were never tested against.

Treat each fallback as a model you have chosen to run in production. Run your evaluation suite against it, validate its responses against the same schemas, and keep a prompt version tuned for it if the primary's prompt does not transfer well. The same discipline applies when a provider retires a model, as covered in what to do when a model is deprecated, and agent evals are what tell you a fallback is safe to use. A self-hosted model can serve as a fallback for some workloads, if you are prepared to run it.

Decide what happens when no model can answer

Sometimes every option is down or too slow. Decide the behaviour for each task type before that happens:

  • Queue and resume: save the task state and continue when a model is available, which suits back-office work with no one waiting.
  • Hand to a person: route the task to a human queue with the context gathered so far, which suits customer-facing work.
  • Refuse clearly: tell the user the agent is unavailable and what to do instead, rather than returning a partial or guessed answer.

Agents that save their progress at each step recover far more cleanly here, the same design that lets them handle a failed tool call without starting again.

Practise the switch before you need it

You cannot wait for a real outage to find out whether failover works. Inject rate-limit errors, timeouts and malformed responses in a staging environment, and confirm that the router switches, the fallback output passes validation, and alerts reach the right people. Monitoring the agent in production should show which model served each request, so you can see a failover happen and spot quality changes after it.

CodeDTX's AI agent maintenance and support and enterprise AI integration work covers this routing, testing and incident planning. To review how your own agents would cope with a provider outage, talk to CodeDTX.

Frequently asked questions

Do we need two LLM providers, or is a second model from one provider enough?

It depends on what you are protecting against. A second model from the same provider helps when one model is overloaded or retired, but it shares the provider's platform, network and account limits. A second provider protects against a wider outage, at the cost of another contract, another data review and another set of prompts to test. Many teams start with one cross-provider fallback for their most important workflows.

Should failover between models be automatic?

Automatic failover suits transient errors and short slowdowns, because a person cannot react fast enough to protect tasks in flight. Longer incidents and quality concerns often need a human decision, since the fallback may cost more or perform differently. A common pattern is an automatic circuit breaker for short failures, plus an operator setting that moves traffic deliberately when an incident runs on or a model change needs approval.

Does a fallback model create data protection or compliance issues?

It can. Sending prompts to a second provider means that provider processes your data, so it needs the same contract terms, data residency review and security assessment as the primary. Check this before an incident, not during one. If a fallback is not approved for certain data, the routing layer should know that and choose another option, such as queueing the task, rather than sending restricted content to an unapproved model.

How do we know a fallback model is good enough?

Run it through the same evaluation suite as the primary model, using realistic tasks from your own workload rather than general benchmarks. Check structured output against your schemas, compare tool-call accuracy and review a sample of answers by hand. Repeat the evaluation when prompts, tools or either model change, because a fallback that passed months ago may no longer match what your agent now asks of it.

Share this post

Contact us to build the right product

Talk to our engineers about your application, the systems it connects to, and what you want to build next.

Get in touch
Two people discussing work with a laptop