Skip to content

Upgrade the AI gateway

Updated 5 min read

The supported deployment target is Envoy AI Gateway 1.1.0, Envoy Gateway 1.8.1, and Gateway API 1.5.1. Use the released upstream Helm charts, including the separate AI Gateway CRD chart. The pack does not install those controllers.

The operator translates the legacy spec.provider.schemaVersion setting into Envoy’s schema.prefix: api/v1 still reaches /api/v1/chat/completions, and the default remains /v1. Bedrock uses its own native paths without a prefix. API schema versions and URL prefixes remain separate in the provider adapter.

Both served and external-provider routes use aigateway.envoyproxy.io/v1beta1. Each endpoint declares its hostname, so /v1/models includes its own catalog without depending on the other endpoint being enabled. Listener references and per-model authentication remain in place. Hostname scoping is not user-level catalog filtering; model invocation still enforces API-key scope and JWT groups. Disabling a provider endpoint removes its route and catalog entries. Its auth policy stays until the generated HTTPRoute is gone, and issued keys are preserved for re-enabling the endpoint.

If the gateway is already on a supported release and only the pack is changing, roll the AI Gateway controller after installing the new operator. This lets it restore cleanup finalizers that older pack operators removed from its resources.

Provider hostnames are normalized to lowercase and have a single trailing DNS root dot removed before gateway resources are generated. The stored spec is not rewritten. Review existing PassthroughModel.spec.provider.hostname values before upgrading: underscores, empty labels, URLs, ports, paths, and whitespace are rejected. Older webhooks admitted some of these values; correct any remaining invalid names before upgrading to avoid reconciliation failures.

helm upgrade does not update files in a chart’s crds/ directory. The pack’s CRDs need an explicit update as well as the gateway CRDs. Otherwise the old PassthroughModel schema still requires a hostname and API-key Secret and does not recognize the new backend and credential fields.

At the pack-operator step of the migration below, run these commands from a checkout of the same release or revision as the chart being installed:

Terminal window
kubectl apply --server-side -f charts/nebari-llm-serving/crds/
kubectl wait --for=condition=Established --timeout=60s \
crd/llmmodels.llm.nebari.dev crd/passthroughmodels.llm.nebari.dev

Then upgrade the Helm release with the matching chart and your existing values. Wait for the updated operator and webhook before applying Bedrock resources. If another tool owns the CRDs, update them through that tool instead; resolve field-ownership conflicts rather than forcing them. Argo CD installations that already manage the CRDs as rendered manifests should keep that existing workflow. Never delete the CRDs to upgrade them: deletion also removes their resources.

Follow the upstream controller upgrade policy and compatibility matrix. Use staged revisions, rather than changing every controller independently:

  1. Record the current chart and operator revisions and inventory all gateway resources, including non-LLM routes on a shared gateway. Check the intervening Envoy Gateway release notes for policy changes. Keep CRDs during rollback; deleting them also deletes their resources.
  2. Install the compatible Gateway API and Envoy Gateway CRDs and upgrade Envoy Gateway to 1.8.1 using the official chart. CRD updates must be explicit: a Helm controller upgrade alone does not reliably update installed CRDs.
  3. Before crossing AI Gateway 0.6, migrate any external OpenAI backends from schema.version to schema.prefix. If the old operator still reconciles these resources, pause that reconciler during the handover. Move any custom AIGatewayRoute.spec.filterConfig into GatewayConfig; this pack does not generate filterConfig.
  4. Upgrade AI Gateway CRDs and controller to 0.7, then update the pack’s CRDs and deploy the new pack operator. Version 0.7 supports hostname-scoped catalogs and the v1beta1 resources the operator now emits. Confirm that all resources reconcile before proceeding to AI Gateway 1.1, using 1.0 as an intermediate controller release. AI Gateway 1.0 has a namespace-discovery bug when the Gateway lives in the Envoy Gateway controller’s namespace: it cannot refresh the sidecar config. In that layout, retain the previous data-plane pod and proceed to 1.1, which deduplicates the namespace lookup.
  5. Roll the Envoy data plane after each AI Gateway controller upgrade so new pods receive its matching ext-proc sidecar. Use an EnvoyProxy pod-template annotation to request the rollout; do not edit generated Deployments. Preserve the ServiceAccount and cloud workload identity. With Argo CD, retain the cert-manager webhook certificate configuration.
  6. Check both inference endpoints, model catalogs, streaming, rejected credentials, local InferencePool models, and non-LLM routes. Expect active connections to reconnect during a data-plane rollout. Only finish the upgrade when the desired GitOps revisions and actual pod images agree.

The v1alpha1 AI Gateway API remains served upstream, but new writes use v1beta1. Reapplying resources migrates their stored representation; do not remove old CRD storage versions until every stored object has been migrated.

For rollback, restore compatible controller/operator revisions through GitOps while retaining the installed CRDs. Returning to AI Gateway 0.5 also requires restoring its path-prefix and catalog configuration; an image-only rollback is not sufficient.

AI Gateway 1.1 makes newer protocol translation, token counting, streaming controls, and tracing available. It does not automatically enable additional providers, failover, or tracing, and does not grant cloud model access. Provider adapters still select the appropriate upstream schema and authentication.

Check generated resources against upstream CRDs

Section titled “Check generated resources against upstream CRDs”

The operator’s TestGatewayAdmission uses a local Kubernetes API server with the released AI Gateway CRDs. It checks admission and stored fields, including path prefixes and hostname scoping. No cluster or cloud credentials are used.

Render ai-gateway-crds-helm version v1.1.0 into a temporary directory, set AI_GATEWAY_CRD_DIR to that directory and KUBEBUILDER_ASSETS to the binaries installed by make setup-envtest, then run from operator:

Terminal window
go test ./internal/controller/reconcilers -run TestGatewayAdmission -v