Download Latest Version agctl-windows-amd64.exe (34.4 MB) Google Add to Preferred Sources
Home / v1.6.0
Name Modified Size InfoDownloads / Week
Parent folder
agctl-darwin-arm64 2026-10-02 32.5 MB
agctl-darwin-arm64.sha256 2026-10-02 85 Bytes
agctl-linux-amd64 2026-10-02 33.6 MB
agctl-linux-amd64.sha256 2026-10-02 84 Bytes
agctl-linux-arm64 2026-10-02 31.3 MB
agctl-linux-arm64.sha256 2026-10-02 84 Bytes
agctl-windows-amd64.exe 2026-10-02 34.4 MB
agctl-windows-amd64.exe.sha256 2026-10-02 90 Bytes
agentgateway-darwin-arm64 2026-10-02 92.2 MB
agentgateway-darwin-arm64.sha256 2026-10-02 100 Bytes
agentgateway-linux-amd64 2026-10-02 97.7 MB
agentgateway-linux-amd64.sha256 2026-10-02 99 Bytes
agentgateway-linux-arm64 2026-10-02 90.5 MB
agentgateway-linux-arm64.sha256 2026-10-02 99 Bytes
agentgateway-windows-amd64.exe 2026-10-02 107.0 MB
agentgateway-windows-amd64.exe.sha256 2026-10-02 105 Bytes
README.md 2026-10-02 54.6 kB
v1.6.0 source code.tar.gz 2026-10-02 7.8 MB
v1.6.0 source code.zip 2026-10-02 9.4 MB
Totals: 19 Items   536.5 MB 1

🎉 Welcome to the 1.6.0 release of the agentgateway project!

This release makes LLM routing more resilient, turns on the Kubernetes AgentgatewayModel API by default, adds zero-configuration cost tracking for common public models, improves OIDC and JWT authentication, and fills important MCP compatibility gaps. It also adds session affinity, per-key local rate limits, graceful draining, and richer OpenTelemetry output.

Before upgrading, review the breaking and behavior changes below. In particular, Anthropic Messages requests now prefer the OpenAI Responses format and provider base URLs without a path now use /.

Artifacts

Docker images are available:

  • cr.agentgateway.dev/agentgateway:v1.6.0
  • cr.agentgateway.dev/controller:v1.6.0

Helm charts are available:

  • cr.agentgateway.dev/charts/agentgateway:v1.6.0
  • cr.agentgateway.dev/charts/agentgateway-crds:v1.6.0
  • cr.agentgateway.dev/charts/agentgateway-standalone:v1.6.0

Binaries are available below.

Quick Start

Follow the Kubernetes or Standalone quick start guide to get started.

🔥 Breaking changes

Anthropic Messages requests now prefer the Responses format

Agentgateway now prefers OpenAI Responses when it converts Anthropic Messages requests. In 1.5, it preferred Chat Completions. This affects OpenAI, Ollama, Groq, Hugging Face, and xAI providers; Azure models other than Claude; and custom providers that advertise both Responses and Chat Completions.

The Responses conversion drops extended-thinking history, while the Chat Completions conversion preserves it. Use the OpenAI provider only for the OpenAI API. For other OpenAI-compatible servers, configure a custom provider and advertise only the formats that the server supports. Advertise only Chat Completions if the server does not implement /v1/responses or clients require extended-thinking history.

As a temporary workaround for a Responses conversion bug, set AGENTGATEWAY_MESSAGES_PREFER_COMPLETIONS=true. This compatibility setting is planned for removal in 1.7.

A built-in model catalog now supplies default request pricing

Requests to common public models now have a cost in logs, traces, metrics, and CEL without a separately configured catalog. As a result, USD API-key budgets in standalone and cost-based policies or CEL expressions in either deployment mode can begin applying to these models after the upgrade.

Review USD budgets and expressions that use llm.cost. Existing catalog configuration can be removed or retained as an overlay to customize the built-in rates.

Provider base URLs without a path now use /

An explicitly configured provider base URL now always sets the complete base path. A URL such as https://api.openai.com therefore uses /. In 1.5, custom and Ollama providers appended /v1, while built-in providers such as OpenAI forwarded the client's path.

Add the provider API path to every configured base URL that currently has no path. For example, change https://api.openai.com to https://api.openai.com/v1 and http://localhost:11434 to http://localhost:11434/v1. URLs that already contain a path and providers that use their default address are unaffected.

Standalone LLM serving paths must match exactly

Standalone's simplified llm configuration now recognizes standard LLM paths only by exact match. A prefixed path such as /tenant-a/v1/messages is passed through without format conversion unless you configure llm.pathPrefix, such as pathPrefix: /tenant-a.

On Kubernetes, model routers attached directly to a Gateway or ListenerSet likewise serve only exact standard LLM paths. Attach models to an HTTPRoute to serve them under a prefix.

The standalone Helm chart requires oidc.enabled

The standalone chart now sets OIDC_COOKIE_SECRET only when the new oidc.enabled value is true; the value defaults to false. If you use oidc.cookieSecretName to protect the UI or another app, set oidc.enabled=true before upgrading or agentgateway will reject the OIDC configuration at startup.

⚠️ Removed

  • Legacy token-count semantics: AGENTGATEWAY_LEGACY_LLM_USAGE_TOKEN_SEMANTICS has been removed. Input and total token counts always include prompt-cache tokens.
  • Server-side defaults in Kubernetes CRDs: CRD schemas no longer persist default values such as action: Allow or tracing.protocol: GRPC. The controller applies the same defaults at runtime, so traffic behavior does not change, but kubectl get -o yaml displays only explicitly configured fields. Update GitOps comparisons and scripts that expected stored defaults.

🔄 Other behavior changes

  • Failover eviction: A backend or virtual model with multiple priority groups now evicts failing targets by default. Defining a health policy replaces this default, so include an eviction policy explicitly when you want to retain failover behavior.
  • Streaming guardrails: A provider guard failure while checking a streamed response or realtime connection now rejects content. Set the guard's failureMode to failOpen to retain the 1.5 behavior.
  • Guardrail results: The guardrails CEL variable includes every guard that ran, including a new allow action. Filter on action != "allow" when you need interventions only.
  • Backend authentication errors: Failures while obtaining provider credentials, such as from OAuth, AWS STS, or Azure, return 502 instead of 500. Local credential application failures return 500 instead of 503.
  • Backend request timeouts: Backend request timeouts now include the time required to read a buffered body, including bodies read by external authorization and CEL response.body expressions. Streamed bodies are unaffected.
  • Policy service timeouts: gRPC external authorization calls now default to 2 seconds; rate limit and external processing calls default to 10 seconds.
  • Access-log levels: Request records use error when agentgateway records a request error and info otherwise. Previously all request records used info.
  • Standalone YAML handling: Unquoted values such as no remain strings instead of becoming booleans. Use true or false for boolean settings. Parse errors now include line numbers, and UI writes preserve comments and formatting where possible.
  • Standalone catalog refresh: The UI's Refresh base costs button downloads the catalog from the repository's main branch. Do not use it with 1.5 or earlier if main contains the 1.6 catalog format.
  • Kubernetes policy conflicts: When equally specific policies set the same field, the oldest policy by metadata.creationTimestamp wins consistently.
  • Kubernetes AI transformations: spec.backend.ai.transformations and spec.backend.ai.finalTransformations reject duplicate field values. Existing policies with duplicates fail validation when reapplied.
  • Kubernetes deployer ownership: The controller no longer overwrites an existing same-named resource that it does not own and reports an error instead.

🌟 New features

LLM routing, formats, and cost controls

  • AgentgatewayModel enabled by default: The Kubernetes agentgatewayModels.enabled Helm value now defaults to true.
  • Wildcard model discovery: /v1/models expands wildcard model configurations into matching model IDs from the catalog. Standalone users can set llm.discovery: disabled to retain the 1.5 behavior of returning the wildcard pattern.
  • More resilient virtual models: Kubernetes virtual models skip broken concrete targets instead of failing the entire model, and backends with one invalid inline policy can remain available with PartiallyValid status.
  • Response idle timeout: A new response idle timeout ends a response when the backend stops sending body data for the configured duration without limiting an active stream's total lifetime.
  • Larger LLM request buffer: The default LLM request buffer increases from 2 MiB to 32 MiB for long-context prompts.
  • Per-page pricing: Model catalog entries support rates.perPage for OCR and document models.
  • Bedrock Mantle routing: Bedrock can choose the Runtime or Mantle endpoint per model, with a configurable endpoint preference.
  • More complete Messages conversion: Conversion between Anthropic Messages and OpenAI formats now carries citations, refusals, strict tool schemas, reasoning effort, and images in tool results. Provider context-overflow errors are translated so Claude Code can compact and retry.
  • Guardrail controls: OpenAI Moderation, Bedrock Guardrails, and Google Model Armor support failureMode. Standalone webhook guards can also configure connection policies, and the built-in credit-card detector now validates the Luhn checksum.

Security and identity

  • OIDC sign-in for standalone: OIDC policies can expose login and logout endpoints. Unauthenticated fetch requests receive 401 instead of an identity-provider redirect, and compressed session cookies accommodate users with more group claims. The UI configures the endpoints automatically.
  • Hashed API keys: The standalone UI stores newly created API keys as hashes by default.
  • Smarter JWT provider selection: When several JWT providers are configured, agentgateway tries providers whose issuer and JWKS key ID match the token. JWT validation also enforces nbf with 60 seconds of clock-skew leeway. Kubernetes policies can configure validation.requiredClaims; standalone MCP authentication no longer requires an audiences field.
  • Cloud backend authentication: AWS assume-role authentication accepts an externalId, and Azure authentication accepts explicit scopes.
  • Network authorization by destination: Network authorization expressions can match destination.address, destination.port, and the TLS SNI value in destination.hostname.
  • Kubernetes CA bundle keys: Backend TLS can read a CA bundle from a Secret or ConfigMap key other than ca.crt, including keys written by trust-manager.

MCP

  • Paginated lists: Tool, prompt, and resource list responses include nextCursor. Federated endpoints return a combined cursor that tracks all targets.
  • MCP data in CEL: Policies can match mcp.methodName, such as tools/call, and standalone access logs can record list results such as mcp.toolsList.
  • Server information overrides: Standalone multiplexed gateways can override reported serverInfo and instructions with an mcp.server block.
  • SSE keep-alive: Standalone can send periodic comments on long-lived MCP streams so idle intermediaries do not close them.
  • Request size enforcement: MCP requests larger than the configured buffer limit return 413.

Traffic management and operations

  • Session affinity on Kubernetes: A CEL expression can derive a session value and consistently send related requests to the same endpoint.
  • Per-key local rate limits: CEL-derived keys give each user, API key, or other value a separate local rate-limit bucket.
  • External processing fail-open: Standalone routes can forward a request when the external processing service fails, including the request body when no body bytes have yet been sent to the processor.
  • Graceful standalone drain: Shutdown can stop accepting new connections at a minimum termination deadline while allowing existing connections to continue until the final termination deadline.
  • Standalone autoscaling and disruption budgets: The standalone Helm chart can create a HorizontalPodAutoscaler and PodDisruptionBudget.
  • Stable ListenerSet ordering: Kubernetes ListenerSets sort by creation time, oldest first, then namespace and name.
  • Reduced Istio permissions: Set istio.enabled=false to remove Istio permissions from the Kubernetes controller ClusterRole.
  • Gateway-aware proxy metrics: The Kubernetes PodMonitor copies the Gateway name pod label onto proxy metrics by default, enabling per-Gateway dashboard filtering.
  • OpenTelemetry updates: OTLP access logs use the agentgateway.access instrumentation scope, and failed requests carry an error_type label on request-duration metrics. Stdout access logs can opt in to OpenTelemetry field names with the OTel access-log preset.

Contributors

Thank you to everyone who contributed code, reviews, documentation, bug reports, and CI improvements for this release!

What's Changed

New Contributors

Full Changelog: https://github.com/agentgateway/agentgateway/compare/v1.5.0...v1.6.0

Source: README.md, updated 2026-10-02