[In preview] Public Preview: Application Gateway for Containers – Inference gateway
- Vendor
- Microsoft
- Product
- Azure Kubernetes Service, Azure Application Gateway
- Type
- Availability
- Announcement date
- 2026-06-24
- Effective date
- Date not published
- Impact
- Affected
In brief
[In preview] Public Preview: Application Gateway for Containers – Inference gateway
What the source says
Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applying OWASP-aligned protections to AI traffic before it reaches your model servers. Improved performance and reliability: Routes around saturated replicas and leverages model server state to lower TTFT, reduce timeouts, and improve GPU utilization. Learn more: Application Gateway for Containers – AI gateway - Inference gateway Application Gateway for Containers - Overview
Read the primary source
Choose your reading
The role changes the reading angle, not the facts, date or level of evidence.
Reading for a salesperson
Who is affected: The audience described by the Microsoft update is represented by: [In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for
Why it matters: The change matters because the official source describes: [In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for
Urgency: Monitor
Next action: Review affected accounts with the customer using the official announcement.
Commercial opportunities
- Use the verified change to open a scoped customer conversation.
Technical actions
- Assess whether the documented change intersects the customer's current stack.
Questions to ask the customer
- Does this documented change affect a product or workload in scope?
Risks and objections
- The official source does not establish facts beyond the quoted material.
Points to confirm
- The effective date is unknown and must be confirmed before scheduling action.
Evidence and traceability
Each excerpt is linked to the primary source and retained for verification.
- Raw capture
2026-08-30T06:15:07.776356+00:00
- Event
[In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applying OWASP-aligned protections to AI traffic before it reaches your model servers. Improved performance and reliability: Routes around saturated replicas and leverages model server state to lower TTFT, reduce timeouts, and improve GPU utilization. Learn more: Application Gateway for Containers – AI gateway - Inference gateway Application Gateway for Containers - Overview
source - Product Azure Kubernetes Service
[In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applying OWASP-aligned protections to AI traffic before it reaches your model servers. Improved performance and reliability: Routes around saturated replicas and leverages model server state to lower TTFT, reduce timeouts, and improve GPU utilization. Learn more: Application Gateway for Containers – AI gateway - Inference gateway Application Gateway for Containers - Overview
source - Product Azure Application Gateway
[In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applying OWASP-aligned protections to AI traffic before it reaches your model servers. Improved performance and reliability: Routes around saturated replicas and leverages model server state to lower TTFT, reduce timeouts, and improve GPU utilization. Learn more: Application Gateway for Containers – AI gateway - Inference gateway Application Gateway for Containers - Overview
source - Insight a salesperson
[In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applyin
source
Back to the public feed