Download Latest Version v2.2.2 source code.zip (64.6 MB)
Email in envelope

Get an email when there's a new version of GPUStack

Home / v2.1.1
Name Modified Size InfoDownloads / Week
Parent folder
gpustack-2.1.1-py3-none-any.whl 2026-03-26 12.6 MB
README.md 2026-03-26 2.9 kB
v2.1.1 source code.tar.gz 2026-03-26 47.2 MB
v2.1.1 source code.zip 2026-03-26 47.7 MB
Totals: 4 Items   107.5 MB 0

Model Catalog Updates

Tuned Qwen3.5 model deployments for optimized inference performance:

  • Qwen3.5-35B-A3B: Achieves +33.0% TPS (when optimized for throughput) or 2.74x generation speed (when optimized for latency) using H200 GPUs. (Report)
  • Qwen3.5-9B: Achieves +15.6% TPS (when optimized for throughput) or 1.26x generation speed (when optimized for latency) using H100 GPUs. (Report)

Enhancements

  • Allow configuring embedded Prometheus and Grafana ports to avoid using default ports. (Issues [#4896])
  • Support authentication using the Anthropic X-API-KEY format. (Issue [#4882])
  • Support configuring route targets with a weight of 0. (Issue [#3772])
  • Inform users to upgrade the driver when using incompatible versions. (Issues [#4873])
  • UI/UX improvements. (Issues [#4833], [#4871])

Bug Fixes

  • Fixed an issue where the worker process repeatedly crashed and restarted. (Issues [#4921], [#4878])
  • Fixed missing required fields for the Ollama model provider. (Issue [#4906])
  • Fixed an issue where user activation and deactivation took ten minutes to take effect. (Issue [#4902])
  • Corrected incorrect context size detection from metadata. (Issue [#4895])
  • Fixed automatic recovery failure for gpustack-worker after a crash in all-in-one deployments. (Issue [#4894])
  • Fixed custom CA certificates not working with OIDC. (Issue [#4893])
  • Fixed an unresponsive worker metrics interface. (Issue [#4879])
  • Resolved a registration failure when a worker name already existed. (Issue [#4875])
  • Fixed token usage not being returned for some streaming responses. (Issue [#4874])
  • Fixed an omission where public MaaS models were not included in the usage statistics list. (Issue [#4864])
  • Fixed a failure to retrieve the model pretrained config during deployment. (Issue [#4855])
  • Fixed an issue where, with NVIDIA vGPU, assigning a UUID value to a backend-visible device blocked workload startup. (Issue [#4844])
  • Resolved an incompatibility where MaxKB used the /v2/rerank endpoint, preventing direct access to GPUStack. (Issue [#4842])
  • Corrected a 404 error when calling the Anthropic API /v1/messages endpoint. (Issue [#4836])
  • Fixed missing arguments for MindIE 2.3.0. (Issue [#4834])
  • Fixed an issue where the Metax C500 GPU could not be detected. (Issue [#4832])
  • Fixed incorrect source handling for legacy custom backends. (Issue [#4827])
  • Fixed a device detection failure in an Ubuntu 24.04 environment with an AMD Radeon RX 7800 XT GPU. (Issue [#4796])

Built-in Inference Backend Updates

New Arrivals

  • CANN: vLLM 0.16.0
  • CUDA: vLLM 0.17.1
  • MACA: vLLM 0.14.0/0.13.0/0.12.0, SGLang 0.5.7
  • ROCm: vLLM 0.17.1
Source: README.md, updated 2026-03-26