[In preview] Public Preview: Application Gateway for Containers – Inference gateway
- Fournisseur
- Microsoft
- Produit
- Azure Kubernetes Service, Azure Application Gateway
- Type
- Disponibilité
- Date annonce
- 2026-06-24
- Date effet
- Date non publiee
- Impact
- Concerné
En bref
[In preview] Public Preview: Application Gateway for Containers – Inference gateway
Ce que dit la source
Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applying OWASP-aligned protections to AI traffic before it reaches your model servers. Improved performance and reliability: Routes around saturated replicas and leverages model server state to lower TTFT, reduce timeouts, and improve GPU utilization. Learn more: Application Gateway for Containers – AI gateway - Inference gateway Application Gateway for Containers - Overview
Lire la source primaire
Choisir votre lecture
Le rôle change l'angle de lecture, pas les faits, la date ni le niveau de preuve.
Lecture pour un commercial
Qui est concerné : The audience described by the Microsoft update is represented by: [In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for
Pourquoi c'est important : The change matters because the official source describes: [In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for
Urgence : Surveiller
Prochaine action : Review affected accounts with the customer using the official announcement.
Opportunités commerciales
- Use the verified change to open a scoped customer conversation.
Actions techniques
- Assess whether the documented change intersects the customer's current stack.
Questions à poser au client
- Does this documented change affect a product or workload in scope?
Risques et objections
- The official source does not establish facts beyond the quoted material.
Points à confirmer
- The effective date is unknown and must be confirmed before scheduling action.
Preuves et traçabilité
Chaque extrait est relié à la source primaire et conservé pour vérification.
- Capture brute
2026-08-30T06:15:07.776356+00:00
- Événement
[In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applying OWASP-aligned protections to AI traffic before it reaches your model servers. Improved performance and reliability: Routes around saturated replicas and leverages model server state to lower TTFT, reduce timeouts, and improve GPU utilization. Learn more: Application Gateway for Containers – AI gateway - Inference gateway Application Gateway for Containers - Overview
source - Produit Azure Kubernetes Service
[In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applying OWASP-aligned protections to AI traffic before it reaches your model servers. Improved performance and reliability: Routes around saturated replicas and leverages model server state to lower TTFT, reduce timeouts, and improve GPU utilization. Learn more: Application Gateway for Containers – AI gateway - Inference gateway Application Gateway for Containers - Overview
source - Produit Azure Application Gateway
[In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applying OWASP-aligned protections to AI traffic before it reaches your model servers. Improved performance and reliability: Routes around saturated replicas and leverages model server state to lower TTFT, reduce timeouts, and improve GPU utilization. Learn more: Application Gateway for Containers – AI gateway - Inference gateway Application Gateway for Containers - Overview
source - Insight un commercial
[In preview] Public Preview: Application Gateway for Containers – Inference gateway Application Gateway for Containers is extending its ingress gateway feature set with new AI gateway capabilities. The new inference gateway capability brings the Kubernetes Gateway API Inference Extension to Application Gateway for Containers, enabling load and model-aware routing for self-hosted generative AI workloads on AKS. Application Gateway for Containers' inference gateway capability is purpose-built for serving LLMs and other inference workloads, routing requests based on model server signals rather than generic load balancing — improving Time to First Token (TTFT) and reducing timeouts under load. Managed Body-Based Router (BBR): A fully managed router that inspects the request body (such as the model field on OpenAI-compatible APIs) to drive model-aware routing — no custom proxy tier required. Bring your own Endpoint Picker (EPP): Choose the EPP implementation that fits your workload to drive metrics-aware endpoint selection across your InferencePool. Secure inference: Pairs natively with the existing web application firewall (WAF) capability in Application Gateway for Containers, applyin
source
Retour au fil public