Multi-region failover for AI APIs is no longer a bespoke project: the three hyperscalers now ship managed routing for model inference, and the engineering work has moved to configuration, policy, and verification. Amazon Bedrock’s global cross-Region inference profiles price approximately 10% below standard regional inference while sending requests to regions you never enabled in your account; Azure OpenAI documents a two-resource pairing pattern whose quota survives the split; and Vertex AI pools capacity across a whole geography with multi-region endpoints. What none of the platforms removes is the decision you still own: which failure you are protecting against, what triggers the switch, and how you prove the secondary path actually works before you need it.
Why Single-Region AI APIs Fail
Regional failures of model APIs are rarely exotic. Three patterns cover most incidents. First, capacity: a single region’s serving pool saturates and your requests come back throttled long before anything is technically down. Second, quota: per-region token limits mean a spike that fits one region’s budget gets clipped when you are pinned to it. Third, the blunt one: control-plane or network incidents that make an entire regional endpoint unreachable for minutes to hours. Each failure mode has a different remedy, and buying more of one does not fix the others. Latency-sensitive designs interact with regional routing too — cold starts on serverless inference get worse, not better, when a failover lands your traffic on a region with no warm capacity for your model. Deciding which failure you are engineering for is the first deliverable; detecting it reliably is an observability problem before it is a routing problem.
Bedrock: Inference Profiles Handle Routing
Bedrock’s cross-Region inference comes in two flavors that look similar and behave differently. Geographic profiles (prefixed eu, us, apac) keep processing inside a boundary — the right default for EU data-residency commitments. Global profiles route to any commercial region and carry the discount noted above. The trap is policy: if an SCP or IAM policy blocks any destination region listed in a geographic profile, routing fails even though your source region is healthy. Global profiles invert the problem — they require allowing “aws:RequestedRegion”: “unspecified”, and conventional region-deny policies silently break them. Verify the destination set with GetInferenceProfile rather than assuming it, because profiles differ by source region.
Observability rides on existing AWS machinery. CloudTrail logs every cross-Region inference request in your source Region, and the additionalEventData.inferenceRegion field tells you where each request actually ran. Without that field in your dashboards you cannot distinguish a routing change from a latency regression, and your failover drills produce no evidence. Wire it into the same pipeline you use for production LLM observability before an incident, not during one. Also note the boundary: inference profiles do not support Provisioned Throughput, so reserved-capacity deployments still need an explicit multi-region plan of their own.
Azure OpenAI: Duplicate, Then Fail Over
Microsoft’s BCDR guidance for Foundry models is refreshingly concrete. Azure OpenAI allocates quota at the subscription-plus-region level, so a primary and a secondary resource can coexist in one subscription without competing for quota. Deploy two resources in different regions, duplicate every model deployment in both, and allocate the full quota to each rather than splitting it — Microsoft states that full allocation yields higher throughput than splitting quota across deployments. Standard deployments with Data Zone or Global routing already spread requests across regions automatically; the manual failover story exists for the harder case, when the primary endpoint itself becomes unreachable.
For provisioned capacity, the recommended topology is a chain: a workload-dedicated PTU deployment overflows to an enterprise PTU pool placed in a different region, and then falls through to Standard deployments. Put the enterprise pool in a different region than your primary Standard deployment so a single regional outage cannot take both tiers down at once. Front the whole thing with a gateway that does circuit breaking — Azure API Management behind Front Door is the documented pairing — because clients should never implement per-provider failover logic themselves.
Vertex AI: Three Endpoint Classes
Google gives you three explicit choices rather than two. Regional endpoints pin processing to one region: lowest latency, clear residency, and full exposure to that region’s capacity and quota. Multi-region endpoints are the middle ground: Claude on Vertex AI answers at aiplatform.us.rep.googleapis.com and aiplatform.eu.rep.googleapis.com, pooling capacity across the US or the EU while keeping processing inside that geography. They carry their own quota pools separate from single-region quotas, and they support prompt caching — the router tries to land a request in the region where its prompt cache already lives, and rebalances within the geography under load. Global endpoints maximize availability: the global endpoint distributes traffic across Google’s internal network to the nearest available regional service, absorbing localized congestion and regional rate limit 429 errors. The price of going global is control — you cannot know or constrain which region processes a request, so residency-bound workloads must stay on regional or multi-region endpoints. Private connectivity adds a wrinkle: Private Google Access is not supported for multi-region endpoints, so private consumers need Private Service Connect instead.
A Failover Pattern That Works
Whatever the platform, the client-side pattern is the same and deliberately boring:
- Pick the failure you are protecting against — capacity throttling, quota exhaustion, or full endpoint loss — and write it down; the routing choice follows from it.
- Prefer provider-managed routing (geographic profiles, Data Zone deployments, multi-region endpoints) as the first line of defense; it fails over without your code knowing.
- Keep a client-side fallback only for endpoint-level loss: a preconfigured secondary endpoint, health-based circuit breaking, and jittered retries for the transient tier.
- Drill quarterly: push production-shaped traffic through the secondary path and record the observed recovery time, not the designed one.
- Instrument routing decisions — inference region, deployment names, endpoint location — so every drill and every incident leaves evidence behind.
| Platform | Managed routing scope | Residency control | Failure you still own |
|---|---|---|---|
| Amazon Bedrock | Geography (eu/us/apac) or global | Geographic profiles keep processing in-boundary | SCP or IAM policy blocking a destination region; Provisioned Throughput excluded from profiles |
| Azure OpenAI | Data Zone / Global Standard spread; manual pairing otherwise | Regional and Data Zone deployment types | Endpoint-level loss; gateway and PTU failover chain configuration |
| Vertex AI | Multi-region (us/eu) or global endpoint | Multi-region keeps processing in-geography | Separate quota pools per endpoint class; private access needs Private Service Connect |
Multi-region failover for AI APIs ends up being mostly policy hygiene and drill discipline. The routers are good; what fails in practice is an SCP that blocks a destination region, a secondary deployment that was never allocated quota, or a team that has never once sent real traffic through the backup path. Spend the effort there.