| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| gpustack-2.1.1-py3-none-any.whl | 2026-03-26 | 12.6 MB | |
| README.md | 2026-03-26 | 2.9 kB | |
| v2.1.1 source code.tar.gz | 2026-03-26 | 47.2 MB | |
| v2.1.1 source code.zip | 2026-03-26 | 47.7 MB | |
| Totals: 4 Items | 107.5 MB | 0 | |
Model Catalog Updates
Tuned Qwen3.5 model deployments for optimized inference performance:
- Qwen3.5-35B-A3B: Achieves +33.0% TPS (when optimized for throughput) or 2.74x generation speed (when optimized for latency) using H200 GPUs. (Report)
- Qwen3.5-9B: Achieves +15.6% TPS (when optimized for throughput) or 1.26x generation speed (when optimized for latency) using H100 GPUs. (Report)
Enhancements
- Allow configuring embedded Prometheus and Grafana ports to avoid using default ports. (Issues [#4896])
- Support authentication using the Anthropic X-API-KEY format. (Issue [#4882])
- Support configuring route targets with a weight of 0. (Issue [#3772])
- Inform users to upgrade the driver when using incompatible versions. (Issues [#4873])
- UI/UX improvements. (Issues [#4833], [#4871])
Bug Fixes
- Fixed an issue where the worker process repeatedly crashed and restarted. (Issues [#4921], [#4878])
- Fixed missing required fields for the Ollama model provider. (Issue [#4906])
- Fixed an issue where user activation and deactivation took ten minutes to take effect. (Issue [#4902])
- Corrected incorrect context size detection from metadata. (Issue [#4895])
- Fixed automatic recovery failure for
gpustack-workerafter a crash in all-in-one deployments. (Issue [#4894]) - Fixed custom CA certificates not working with OIDC. (Issue [#4893])
- Fixed an unresponsive worker metrics interface. (Issue [#4879])
- Resolved a registration failure when a worker name already existed. (Issue [#4875])
- Fixed token usage not being returned for some streaming responses. (Issue [#4874])
- Fixed an omission where public MaaS models were not included in the usage statistics list. (Issue [#4864])
- Fixed a failure to retrieve the model pretrained config during deployment. (Issue [#4855])
- Fixed an issue where, with NVIDIA vGPU, assigning a UUID value to a backend-visible device blocked workload startup. (Issue [#4844])
- Resolved an incompatibility where MaxKB used the
/v2/rerankendpoint, preventing direct access to GPUStack. (Issue [#4842]) - Corrected a 404 error when calling the Anthropic API
/v1/messagesendpoint. (Issue [#4836]) - Fixed missing arguments for MindIE 2.3.0. (Issue [#4834])
- Fixed an issue where the Metax C500 GPU could not be detected. (Issue [#4832])
- Fixed incorrect source handling for legacy custom backends. (Issue [#4827])
- Fixed a device detection failure in an Ubuntu 24.04 environment with an AMD Radeon RX 7800 XT GPU. (Issue [#4796])
Built-in Inference Backend Updates
New Arrivals
- CANN:
vLLM 0.16.0 - CUDA:
vLLM 0.17.1 - MACA:
vLLM 0.14.0/0.13.0/0.12.0,SGLang 0.5.7 - ROCm:
vLLM 0.17.1