| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| gpustack-chart-2.2.2.tgz | 2026-07-24 | 53.5 kB | |
| gpustack-2.2.2-py3-none-any.whl | 2026-07-24 | 18.9 MB | |
| README.md | 2026-07-24 | 9.5 kB | |
| v2.2.2 source code.tar.gz | 2026-07-24 | 64.0 MB | |
| v2.2.2 source code.zip | 2026-07-24 | 64.6 MB | |
| Totals: 5 Items | 147.6 MB | 3 | |
đ Security
This release fixes two security vulnerabilities. Users are strongly advised to upgrade immediately.
- Missing authorization on several cluster and resource endpoints, allowing cross-user access and cluster takeover.
- Affected versions: v2.2.0âv2.2.1
- Fixed in: v2.2.2
- CVE assignment is pending; this release note will be updated when the identifier becomes available.
- SAML callback trusts an unverified
SAMLResponse, allowing authentication bypass and account takeover. - Affected versions: v0.7.1âv2.2.1
- Fixed in: v2.2.2
- CVE assignment is pending; this release note will be updated when the identifier becomes available.
đ Model Catalog Updates
LLM: GLM-5.2, Tencent Hy3, MiniMax-M3, moonshotai/Kimi-K2.7-Code, google/diffusiongemma-26B-A4B-it.
⨠Enhancements
- Allow overriding CUDA/driver compatibility filtering for inference engine versions (#5674).
- Support the
{{gpu_count}}placeholder in custom backend arguments for dynamictensor-parallel-size(#5510). - Improve manual GPU selection UX during model deployment (#5671).
- Enhanced GPU Service metering (#5716).
- Show idle storage in usage (#5697).
- Helm: add
extraVolumes/extraVolumeMountsto the server chart (mount private CA or extra certs via values) (#5848). - Rename the worker Service from
workertogpustack-workerfor clarity (#5891). - Inform users of supported Kubernetes versions (#5772).
- Add an upgrade guide for Helm installation (#5764).
- Add Higress-related dashboard for embedded gateway mode (#5717).
- Add an option to skip TLS certificate verification for SSO providers (CAS/OIDC, etc.) (#5830).
- Rename
GPUSTACK_USAGE_ROLLUP_TIMEZONEto a platform-wideGPUSTACK_TIMEZONE(#5900). - Improve legend colors in usage charts for better distinguishability (#5691).
đ Bug Fixes
- Fixed multi-node DEP mode passing incompatible args when running GLM 5.1 (#5538).
- Fixed manual multi-node DEP deployment blocked by the parallel-size check (#5539).
- Fixed MiniMax M2.7 staying Pending on 910C with 16 GPUs when using CP parallelism (#5612).
- Fixed distributed deployment TP validation issue on DGX Spark GB10 (#5779).
- Fixed intermittent "address already in use" error when starting distributed vLLM MP deployments (#5657).
- Fixed route status incorrectly reverting to unavailable after the browser loses focus while modifying weight (#5608).
- Fixed
/v2/modelsfailure (#5734). - Fixed failure to generate audio after switching models (#5649).
- Fixed benchmark job failing on Kimi-K2.6 due to a missing
tiktokendependency (#5689). - Fixed single-node deployment logs mixed with irrelevant ray-head logs (#5696).
- Fixed log writer stopping unexpectedly, leaving the UI with only partial logs (#5740).
- Fixed being able to log in with a deleted SSH key still bound to a GPU instance (#5640).
- Fixed regular users failing to create a GPU instance with a persistent volume (#5692).
- Fixed GPU instance usage not being recorded correctly (#5710).
- Fixed PVC not being cleaned up after storage deletion (#5802).
- Fixed wrong instance type for GPU in usage (#5681).
- Fixed wrong usage legend when grouping GPU instance usage by instance type (#5700).
- Fixed missing data in usage trends caused by incorrect pagination parameters (#5690).
- Fixed Models Used = 0 on the Usage Tokens page when a deleted model is selected (#5795).
- Fixed duplicate 'Deleted' labels for deleted models in the model filter dropdown (#5798).
- Fixed regular users unable to view token usage (#5687).
- Fixed missing data in GPU instance and storage exports (#5833).
- Fixed API keys with Platform Management & Model Access permissions unable to retrieve a newly deployed model instance via
/v2/model-instances(#5686). - Fixed an admin-generated API key unable to query resources manually created by the admin in the UI (#5701).
- Fixed the GPUStack Operator pod exiting with an error when deployed via Helm on Kubernetes v1.35 (#5698).
- Fixed worker port config (
worker_port/worker_metrics_port) not applied to the generated Service (#5846). - Fixed the ext-auth WasmPlugin applying globally, authenticating unrelated traffic when the Higress gateway is shared (#5744).
- Fixed
GPUSTACK_GRAFANA_URLusing the GPUStack URL protocol instead of the configured HTTP protocol (#5762). - Fixed the worker exposing
/serveLogsand/debugwithout authentication (#5836). - Fixed
reload-configon a worker trying to apply server debug config and failing with 401 (#5867). - Fixed server performance degradation (#5892).
- Fixed high disk usage from PostgreSQL logs (#5774).
- Fixed a
gpustack-workerleak of/dev/dri/card1file descriptors and anonymous memory on hybrid AMD iGPU + NVIDIA systems (#5342). - Fixed deleting a model provider dropping other providers' McpBridge registries (prefix collision on
provider-<id>) (#5927). - Fixed the custom model library doubling the model count on each refresh (#5685).
- Fixed worker dashboard only showing CPU and memory for a single node (#5759).
- Fixed worker version shown in the UI not matching the installed version (#5776).
- Fixed UI issues when switching cluster type between Model Service and GPU Service (#5799).
- Fixed Instance List "Filter by cluster" not taking effect (#5702, #5823).
- Fixed the user filter not sending a filtered request (#5723).
- Fixed the user list "Name" column sort throwing a backend error and clearing the list (#5706).
- Fixed the Providers list sort button not working (#5705).
- Fixed empty-state hint vertical alignment (#5860).
đ§ Built-in Inference Backend Updates
New Arrival
- CANN 9.0
SGLang 0.5.15.post1 / 0.5.14(910B / A3)- CUDA 13.0
vLLM 0.25.1 / 0.24.0SGLang 0.5.15.post1 / 0.5.14- CUDA 12.9
vLLM 0.25.1 / 0.24.0SGLang 0.5.15.post1 / 0.5.14- DTK 26.04
vLLM 0.18.1SGLang 0.5.10- HGGC 13.0
vLLM 0.23.0 / 0.20.1 / 0.19.0 / 0.18.0SGLang 0.5.12 / 0.5.10 / 0.5.9- MACA 3.7
vLLM 0.21.0 / 0.20.0SGLang 0.5.11 / 0.5.10- MACA 3.5
SGLang 0.5.9- ROCm 7.2
vLLM 0.25.1 / 0.24.0SGLang 0.5.15.post1 / 0.5.14