Download Latest Version v2.2.2 source code.zip (64.6 MB)
Email in envelope

Get an email when there's a new version of GPUStack

Home / v2.2.2
Name Modified Size InfoDownloads / Week
Parent folder
gpustack-chart-2.2.2.tgz 2026-07-24 53.5 kB
gpustack-2.2.2-py3-none-any.whl 2026-07-24 18.9 MB
README.md 2026-07-24 9.5 kB
v2.2.2 source code.tar.gz 2026-07-24 64.0 MB
v2.2.2 source code.zip 2026-07-24 64.6 MB
Totals: 5 Items   147.6 MB 3

🔒 Security

This release fixes two security vulnerabilities. Users are strongly advised to upgrade immediately.

  • Missing authorization on several cluster and resource endpoints, allowing cross-user access and cluster takeover.
  • Affected versions: v2.2.0–v2.2.1
  • Fixed in: v2.2.2
  • CVE assignment is pending; this release note will be updated when the identifier becomes available.
  • SAML callback trusts an unverified SAMLResponse, allowing authentication bypass and account takeover.
  • Affected versions: v0.7.1–v2.2.1
  • Fixed in: v2.2.2
  • CVE assignment is pending; this release note will be updated when the identifier becomes available.

🚀 Model Catalog Updates

LLM: GLM-5.2, Tencent Hy3, MiniMax-M3, moonshotai/Kimi-K2.7-Code, google/diffusiongemma-26B-A4B-it.

✨ Enhancements

  • Allow overriding CUDA/driver compatibility filtering for inference engine versions (#5674).
  • Support the {{gpu_count}} placeholder in custom backend arguments for dynamic tensor-parallel-size (#5510).
  • Improve manual GPU selection UX during model deployment (#5671).
  • Enhanced GPU Service metering (#5716).
  • Show idle storage in usage (#5697).
  • Helm: add extraVolumes/extraVolumeMounts to the server chart (mount private CA or extra certs via values) (#5848).
  • Rename the worker Service from worker to gpustack-worker for clarity (#5891).
  • Inform users of supported Kubernetes versions (#5772).
  • Add an upgrade guide for Helm installation (#5764).
  • Add Higress-related dashboard for embedded gateway mode (#5717).
  • Add an option to skip TLS certificate verification for SSO providers (CAS/OIDC, etc.) (#5830).
  • Rename GPUSTACK_USAGE_ROLLUP_TIMEZONE to a platform-wide GPUSTACK_TIMEZONE (#5900).
  • Improve legend colors in usage charts for better distinguishability (#5691).

🐛 Bug Fixes

  • Fixed multi-node DEP mode passing incompatible args when running GLM 5.1 (#5538).
  • Fixed manual multi-node DEP deployment blocked by the parallel-size check (#5539).
  • Fixed MiniMax M2.7 staying Pending on 910C with 16 GPUs when using CP parallelism (#5612).
  • Fixed distributed deployment TP validation issue on DGX Spark GB10 (#5779).
  • Fixed intermittent "address already in use" error when starting distributed vLLM MP deployments (#5657).
  • Fixed route status incorrectly reverting to unavailable after the browser loses focus while modifying weight (#5608).
  • Fixed /v2/models failure (#5734).
  • Fixed failure to generate audio after switching models (#5649).
  • Fixed benchmark job failing on Kimi-K2.6 due to a missing tiktoken dependency (#5689).
  • Fixed single-node deployment logs mixed with irrelevant ray-head logs (#5696).
  • Fixed log writer stopping unexpectedly, leaving the UI with only partial logs (#5740).
  • Fixed being able to log in with a deleted SSH key still bound to a GPU instance (#5640).
  • Fixed regular users failing to create a GPU instance with a persistent volume (#5692).
  • Fixed GPU instance usage not being recorded correctly (#5710).
  • Fixed PVC not being cleaned up after storage deletion (#5802).
  • Fixed wrong instance type for GPU in usage (#5681).
  • Fixed wrong usage legend when grouping GPU instance usage by instance type (#5700).
  • Fixed missing data in usage trends caused by incorrect pagination parameters (#5690).
  • Fixed Models Used = 0 on the Usage Tokens page when a deleted model is selected (#5795).
  • Fixed duplicate 'Deleted' labels for deleted models in the model filter dropdown (#5798).
  • Fixed regular users unable to view token usage (#5687).
  • Fixed missing data in GPU instance and storage exports (#5833).
  • Fixed API keys with Platform Management & Model Access permissions unable to retrieve a newly deployed model instance via /v2/model-instances (#5686).
  • Fixed an admin-generated API key unable to query resources manually created by the admin in the UI (#5701).
  • Fixed the GPUStack Operator pod exiting with an error when deployed via Helm on Kubernetes v1.35 (#5698).
  • Fixed worker port config (worker_port/worker_metrics_port) not applied to the generated Service (#5846).
  • Fixed the ext-auth WasmPlugin applying globally, authenticating unrelated traffic when the Higress gateway is shared (#5744).
  • Fixed GPUSTACK_GRAFANA_URL using the GPUStack URL protocol instead of the configured HTTP protocol (#5762).
  • Fixed the worker exposing /serveLogs and /debug without authentication (#5836).
  • Fixed reload-config on a worker trying to apply server debug config and failing with 401 (#5867).
  • Fixed server performance degradation (#5892).
  • Fixed high disk usage from PostgreSQL logs (#5774).
  • Fixed a gpustack-worker leak of /dev/dri/card1 file descriptors and anonymous memory on hybrid AMD iGPU + NVIDIA systems (#5342).
  • Fixed deleting a model provider dropping other providers' McpBridge registries (prefix collision on provider-<id>) (#5927).
  • Fixed the custom model library doubling the model count on each refresh (#5685).
  • Fixed worker dashboard only showing CPU and memory for a single node (#5759).
  • Fixed worker version shown in the UI not matching the installed version (#5776).
  • Fixed UI issues when switching cluster type between Model Service and GPU Service (#5799).
  • Fixed Instance List "Filter by cluster" not taking effect (#5702, #5823).
  • Fixed the user filter not sending a filtered request (#5723).
  • Fixed the user list "Name" column sort throwing a backend error and clearing the list (#5706).
  • Fixed the Providers list sort button not working (#5705).
  • Fixed empty-state hint vertical alignment (#5860).

🔧 Built-in Inference Backend Updates

New Arrival

  • CANN 9.0
  • SGLang 0.5.15.post1 / 0.5.14 (910B / A3)
  • CUDA 13.0
  • vLLM 0.25.1 / 0.24.0
  • SGLang 0.5.15.post1 / 0.5.14
  • CUDA 12.9
  • vLLM 0.25.1 / 0.24.0
  • SGLang 0.5.15.post1 / 0.5.14
  • DTK 26.04
  • vLLM 0.18.1
  • SGLang 0.5.10
  • HGGC 13.0
  • vLLM 0.23.0 / 0.20.1 / 0.19.0 / 0.18.0
  • SGLang 0.5.12 / 0.5.10 / 0.5.9
  • MACA 3.7
  • vLLM 0.21.0 / 0.20.0
  • SGLang 0.5.11 / 0.5.10
  • MACA 3.5
  • SGLang 0.5.9
  • ROCm 7.2
  • vLLM 0.25.1 / 0.24.0
  • SGLang 0.5.15.post1 / 0.5.14
Source: README.md, updated 2026-07-24