Four Ways to Reach a Model in Another Azure Region From Microsoft Foundry
Author(s): Dave R | Microsoft Azure & AI MVP ☁️ Originally published on Towards AI.

- Author(s): Dave R | Microsoft Azure & AI MVP ☁️ Originally published on Towards AI.
- Microsoft Foundry model availability is regional, so the model or Foundry Agent Service feature you need can live outside the region your project was approved for.
- Pattern 1 is a direct Foundry-to-Foundry connection that provides static model governance but cannot enforce per-call policies like token limits, metrics, caching, or failover.
Latest Machine Learning
Four Ways to Reach a Model in Another Azure Region From Microsoft Foundry
Dave R - Microsoft Azure & AI MVP☁️
44 likes
September 25, 2026
Author(s): Dave R | Microsoft Azure & AI MVP ☁️
Originally published on Towards AI .
How each pattern handles identity, routing, and private networking, and which Foundry features stop working when a gateway sits in the path.
Microsoft Foundry model availability is regional, so the model or Foundry Agent Service feature you need can live outside the region your project was approved for. Foundry supports four ways to reach it, and they differ in who owns identity, routing, and the network path. One of them also makes first-party tools such as SharePoint grounding fail with bad_request . This guide compares all four patterns, shows the API Management option running on a fully private network, lists which features survive the hop, and ends with a decision flow you can apply to your own landing zone.
Four patterns connecting a Foundry project in one Azure region to model deployments in another region.
The article explains why cross-region model access is an ownership decision across three boundaries—control plane, identity, and the data path—and then details four supported patterns. Pattern 1 is a direct Foundry-to-Foundry connection that provides static model governance but cannot enforce per-call policies like token limits, metrics, caching, or failover. Pattern 2 uses Azure API Management (APIM) as a model gateway, with one parameterized route per deployment type so a platform team can centrally apply managed identity, routing, and response labeling while gaining token budgets, metrics, caching, and resilience controls; it also shows how to run this fully privately using private endpoints and private DNS.
Pattern 3 places APIM on the agent ingress for governance and observability, but it cannot replace the caller’s identity on the agent surface, so key first-party on-behalf-of capabilities depend on Foundry validating the caller token. Pattern 4 adapts routing for the Responses API, where the model name is in the request body; it relies on Foundry’s dynamic model connections rather than a single APIM route, trading centralized routing visibility for cleaner dispatch. Finally, it summarizes what still works when traffic goes through a gateway (e.g., state-bearing agents and most routing) and what breaks or must be moved (notably first-party grounding with gateway-routed calls), provides guidance for private networking end-to-end, and closes with a decision flow and recommended build order.
Read the full blog for free on Medium .
Join thousands of data leaders on the AI newsletter . Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup , an AI-related product, or a service, we invite you to consider becoming a sponsor .
Published via Towards AI
Towards AI - Medium
Towards AI Academy
We Build Enterprise-Grade AI. We'll Teach You to Master It Too.
15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.
Start free — no commitment:
→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day
→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages
Our courses:
→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.
→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.
→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.
Note: Article content contains the views of the contributing authors and not Towards AI.
Sources
Related stories

Microsoft expands Azure AI and HPC infrastructure with AMD
AI workloads are scaling faster than any single infrastructure approach can support — with more models, new agent-driven workloads and surging compute demand driving the need for greater specialization across the stack. To meet this need, Microsoft continues to evolve Azure’s.

Red Hat Shrinks Nemotron 3.5 Lightning's 30B Agent Model by Half With FP8
Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B. Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.

NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI
Local AI is becoming more useful by the token. As AI agents move from experiments into everyday development, increasingly capable open models are shrinking to fit on more devices, giving builders more to run locally.

What Is Jev? A Guide to TypeSafe AI’s System One Model
What Is Jev? A Guide to TypeSafe AI’s System One Model Agents run in a loop: an LLM decides what to do, a tool executes, a model evaluates the results, and then continues in that loop until the task is complete.

The most Viral new AI model isn’t an LLM at all
Jev by Typesafe AI. Good Morning I don’t usually nerd out on non-LLM machine learning announcements. But extraordinary claims have been made.

As AI Grows More Complex, Model Builders Rely on NVIDIA
Unveiling what it describes as the most capable model series yet for professional knowledge work, OpenAI launched GPT-5.2 in December.

Vulnerabilities in AI Agents Expose Trust Gap Flaw
Google and other organizations have acknowledged vulnerabilities in their AI agents, which exploit trust gaps in the Model Context Protocol.

China Telecom's Xing4.0 Runs a 29B Coding Agent on 19GB Locally
Subtopic Mixture Of Experts · Long Context · Vision Language China Telecom released Xing4.0-29B-A4B , a 29B MoE with only 4B active parameters. Native 256K context extensible to 512K using MLA attention and 64 routed experts.