[In preview] Public Preview: Application Gateway for Containers – Inference gateway

Vendor
Microsoft
Product
Azure Kubernetes Service, Azure Application Gateway
Type
Availability
Announcement date
2026-06-24
Effective date
Date not published
Impact
Affected

In brief

[In preview] Public Preview: Application Gateway for Containers – Inference gateway

What the source says

Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applying OWASP-aligned protections to AI traffic before it reaches your model servers. Improved performance and reliability: Routes around saturated replicas and leverages model server state to lower TTFT, reduce timeouts, and improve GPU utilization. Learn more: Application Gateway for Containers – AI gateway - Inference gateway Application Gateway for Containers - Overview

Continue from here

Read the primary source

Choose your reading

The role changes the reading angle, not the facts, date or level of evidence.

Reading for a salesperson

Who is affected: The audience described by the Microsoft update is represented by: [In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for

Why it matters: The change matters because the official source describes: [In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for

Urgency: Monitor

Next action: Review affected accounts with the customer using the official announcement.

Commercial opportunities

  • Use the verified change to open a scoped customer conversation.

Technical actions

  • Assess whether the documented change intersects the customer's current stack.

Questions to ask the customer

  • Does this documented change affect a product or workload in scope?

Risks and objections

  • The official source does not establish facts beyond the quoted material.

Points to confirm

  • The effective date is unknown and must be confirmed before scheduling action.

Evidence and traceability

Each excerpt is linked to the primary source and retained for verification.

  • Raw capture
    2026-08-30T06:15:07.776356+00:00
  • Event
    [In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applying OWASP-aligned protections to AI traffic before it reaches your model servers. Improved performance and reliability: Routes around saturated replicas and leverages model server state to lower TTFT, reduce timeouts, and improve GPU utilization. Learn more: Application Gateway for Containers – AI gateway - Inference gateway Application Gateway for Containers - Overview
    source
  • Product Azure Kubernetes Service
    [In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applying OWASP-aligned protections to AI traffic before it reaches your model servers. Improved performance and reliability: Routes around saturated replicas and leverages model server state to lower TTFT, reduce timeouts, and improve GPU utilization. Learn more: Application Gateway for Containers – AI gateway - Inference gateway Application Gateway for Containers - Overview
    source
  • Product Azure Application Gateway
    [In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applying OWASP-aligned protections to AI traffic before it reaches your model servers. Improved performance and reliability: Routes around saturated replicas and leverages model server state to lower TTFT, reduce timeouts, and improve GPU utilization. Learn more: Application Gateway for Containers – AI gateway - Inference gateway Application Gateway for Containers - Overview
    source
  • Insight a salesperson
    [In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applyin
    source

Back to the public feed