All articles
AI & LLMsSeptember 25, 2026 7 min read

Fallback Chains for LLM Calls: Designing Around Provider Outages Without Breaking Latency SLAs

When Anthropic goes down for two hours, your product still needs to answer. Here's how we design fallback chains across Claude, GPT, and Gemini without blowing latency budgets or shipping garbage responses.

Fallback Chains for LLM Calls: Designing Around Provider Outages Without Breaking Latency SLAs

The last time a major LLM provider had a multi-hour incident, half of our clients' support bots went silent and the other half kept humming. The difference wasn't luck — it was whether the team had built a real fallback chain or just written try/except around a single API call. If your product's core loop depends on one model from one vendor, you're one status-page red bar away from an incident channel full of angry PMs.

This is a walkthrough of how we design multi-provider fallback chains for LLM calls in production: what to fall back to, when to give up, and how to keep quality from silently degrading.

Why single-provider setups keep breaking

Every major provider has had extended incidents in the last two years. Anthropic, OpenAI, and Google all publish status pages, and if you actually read the incident history rather than the uptime marketing, you'll see multi-hour degradations happen a few times per quarter per vendor. That's before you count rate-limit spikes, regional issues, and the classic "model is available but P99 latency just went from 4s to 40s" scenario.

The naive response is "just retry." But retries against the same endpoint during an incident do nothing except waste your token budget and make the outage worse for everyone. A real fallback strategy has to answer four questions:

  1. What signals trigger a failover?
  2. Which alternate model handles the request?
  3. How do we keep the prompt portable across providers?
  4. How do we know the fallback response is still acceptable?

Designing the trigger conditions

Not every error should trigger a fallback. If the user sent malformed input and the model returns a 400, falling over to a second provider will just return another 400 and double your cost. We split failure modes into three buckets:

  • Retry on same provider: transient 5xx, timeouts under the latency budget, overloaded_error from Anthropic (see Anthropic's error docs), 429s with a retry-after under 2 seconds.
  • Fail over to next provider: repeated 5xx after 1 retry, timeouts exceeding the latency budget, 529 overloaded responses, connection errors, or 429s with long backoff.
  • Return error to caller: 400s, content policy violations, authentication failures, and anything that indicates the request itself is the problem.

The key is treating latency as a failure signal, not just error codes. If Claude Sonnet is technically responding but taking 25 seconds for a request that normally takes 3, that's an outage from the user's perspective.

A concrete budget

We usually give each provider attempt a hard timeout equal to p95_latency * 2.5 for that model on that prompt shape. If the first provider blows through it, we cancel the request and move on. The total request budget is roughly sum(provider_timeouts) + small buffer, and we tell product teams to design UX around the total, not the happy path.

Picking the fallback order

The fallback order isn't just "whichever model is second-best." It depends on the task. Here's how we usually think about it:

  • Structured extraction / tool calls: Claude Sonnet → GPT-4 class → Gemini 2.x. All three support structured outputs and tool use, but prompt behaviour differs enough that we validate each with the same eval set.
  • Long-context summarisation: Gemini → Claude → GPT. Gemini's context window and pricing on long inputs tend to win, but Claude is a strong second.
  • Cheap high-volume classification: a small model on one provider, falling back to a small model on another. Never fall back from a small cheap model to a frontier model unless you've priced the incident scenario — a stuck fallback loop on GPT-4 class pricing can burn a month's budget in an afternoon.

One rule we enforce: never make the fallback more expensive than 3x the primary. If your primary is a $0.25/M input token model and the fallback is $15/M, an incident becomes a financial event, not just a reliability event.

Keeping prompts portable

This is where most fallback implementations quietly fail. Teams write a prompt tuned for Claude's XML-tag style, then fall back to GPT and get worse output that technically parses but is subtly wrong. Users don't get an error — they get a bad answer.

We use a thin abstraction layer that:

  • Stores prompts as structured templates (system, user, examples, output schema) rather than raw strings
  • Renders per-provider variants at call time
  • Runs the same eval set against every provider in the chain before deployment

Here's a stripped-down version of the router pattern:

from dataclasses import dataclass
from typing import Callable
import asyncio

@dataclass
class ProviderAttempt:
    name: str
    call: Callable
    timeout_s: float
    max_retries: int = 1

async def run_with_fallback(prompt, chain: list[ProviderAttempt]):
    last_error = None
    for attempt in chain:
        for retry in range(attempt.max_retries + 1):
            try:
                return await asyncio.wait_for(
                    attempt.call(prompt),
                    timeout=attempt.timeout_s,
                )
            except (TimeoutError, ProviderOverloadedError) as e:
                last_error = e
                if retry < attempt.max_retries:
                    await asyncio.sleep(0.5 * (retry + 1))
                    continue
                break  # move to next provider
            except NonRetryableError:
                raise  # 400s, auth, policy violations
    raise AllProvidersFailed(last_error)

This is deliberately boring. The interesting logic lives in the provider-specific call functions and in the eval suite that gates prompt changes.

Guarding against silent quality drops

Every fallback response gets tagged with the provider that produced it. That tag flows into logs, into user feedback records, and into your eval dashboards. If Gemini responses have a 3x thumbs-down rate compared to Claude for the same task, you want to know that within a day, not next quarter. We've written about this pattern more in our engineering blog.

Cost guardrails during incidents

An incident is exactly when your traffic spikes (users retry, your own retries stack up) and your cheapest provider is unavailable. This is the worst possible time to have loose cost controls.

Things we bake in from day one:

  • Per-provider daily spend caps. If GPT-4 class usage exceeds N dollars in a rolling 24h window, the fallback stops using it and returns a degraded response instead.
  • Circuit breakers per provider. If a provider fails more than 50% of requests in a 60-second window, we stop trying it entirely for the next 5 minutes. Constantly retrying a dead provider adds latency for zero benefit.
  • Request coalescing on identical inputs. During an incident, users mash retry buttons. Deduplicate at the edge.

What "degraded mode" actually looks like

Sometimes every model in your chain is unhappy, or your cost cap has fired. You need a defined degraded mode, not a 500 error. Options we've shipped:

  • Serve a cached response if the input is a near-duplicate of something answered recently
  • Fall back to a rule-based or retrieval-only response with a UI note ("answering from indexed docs while our AI service recovers")
  • Queue the request and email the user when the answer is ready
  • Return a clear error with an ETA, not a spinner that eventually times out

The worst outcome is silent degradation where the user thinks they got a real answer but didn't. Any degraded mode should be observable to the user, either through a UI indicator or a different response format.

Testing the chain before you need it

Gameday your fallback chain. Once a quarter, we deliberately block the primary provider at the network level in a staging environment mirrored to production traffic patterns and watch what happens. The first time we did this we discovered:

  • Our timeout was 60 seconds, so the UI hung for a minute before failing over
  • Gemini's JSON mode produced subtly different key ordering that broke a downstream parser
  • The circuit breaker was per-instance, not per-cluster, so every pod had to independently discover the outage

None of that would have shown up in unit tests. You find it by breaking things on purpose.

Where we'd start

If you have a single-provider LLM call in production today, don't try to build the whole thing at once. In order:

  1. Add structured logging for latency, error type, and provider on every call. You cannot design a fallback without knowing your current failure profile.
  2. Wrap the call in a timeout tighter than your UX budget, and a circuit breaker per provider.
  3. Port your top three prompts to a second provider and run your eval set against both. If the eval scores are close, you have a viable fallback. If they're not, fix the prompt before fixing the reliability.
  4. Ship the fallback behind a feature flag, force it on for 1% of traffic for a week, and watch quality metrics.
  5. Only then wire in a third provider and a defined degraded mode.

Reliability for LLM features is not a one-line config change. It's the same discipline you'd apply to any dependency you don't control — you just have to apply it to three of them at once.

#LLMs#Reliability#Architecture#Claude#OpenAI#Gemini

Want a team like ours?

72Technologies builds production software for the kind of teams who actually read this blog.

Start a project