Download Latest Version v0.9 source code.zip (4.7 MB) Google Add to Preferred Sources
Home / v0.9
Name Modified Size InfoDownloads / Week
Parent folder
README.md 2026-09-16 8.4 kB
v0.9 source code.tar.gz 2026-09-16 3.5 MB
v0.9 source code.zip 2026-09-16 4.7 MB
Totals: 3 Items   8.3 MB 0

Oumi v0.9 Release Notes

This release makes Oumi an end-to-end stack for agentic, tool-using models: you can now train on tool-calling data, execute tool calls against real environments (databases, HTTP endpoints, deterministic lookups, and model-simulated tools), synthesize multi-turn tool conversations, and run GRPO/RL over any of those environments with verl. It also adds a multi-criteria RubricJudge, partial-failure-tolerant inference/synthesis/judging, a Modal GPU launcher and SkyPilot-routed Slurm, and new Gemma 4 recipes.

Highlights

🛠️ Tool-use training

Train models to call tools end-to-end:

  • Tool-aware chat templates — access the tool-call section, auto-detect the tool-call form a template expects (#2605, [#2603], [#2604])
  • Trainable tool adapters (#2658)
  • Collator masking that correctly remasks tool-response spans, including responses nested in an assistant turn (#2594)
  • Tool-calling SFT and text DPO over JSONL, with serialized tool arguments decoded during tokenization (#2642, [#2517])
  • Canonical tool-call Conversation shape + JSON-string ↔ dict argument helpers (#2602, [#2601])

🌐 Tool-execution Environments

The new oumi.environments subsystem supports executing tools calls with various backends.

  • ExecutableEnvironment / ExecutableTool — the base abstraction for tools that actually run. (#2527)
  • DatabaseExecutableEnvironment — run tool calls against a database session, with rollback-isolated teardown (see the EHR example). (#2528, [#2545])
  • EndpointEnvironment — dispatch tool calls to HTTP endpoints. (#2625)
  • Lookup (deterministic) and simulated environments — deterministic lookup tables validated against tool schemas, with argument defaults applied and any JSON value as output; model-simulated tool execution for synthesis. (#2623, [#2536], [#2573], [#2574])
  • Tool executors are resolvable by registry name, so environments can wire tools declaratively. (#2555)

🎯 OumiVerlTool — RL over any Oumi environment

OumiVerlTool bridges Oumi's executable environments into verl, so you can run verl GRPO tool-use RL over any Oumi environment. Custom judge rewards are now supported in verl GRPO, multi-turn prompts are supported in the GRPO data path, conversation metadata is preserved through GRPO datasets, and per-run verl metrics are mirrored to verl_metrics.jsonl with richer logging. (#2569, [#2624], [#2542], [#2644], [#2643], [#2646])

⚖️ RubricJudge — multi-criteria evaluation

A new RubricJudge scores a response against several named criteria in a single judge call, instead of one boolean per pass — useful both for evaluation and as a GRPO reward. Ships with example configs under configs/projects/judges/rubric and documentation. (#2626, [#2628], [#2629])

🧬 Conversation synthesis

Synthesize multi-turn tool-use conversations:

  • Stateful and stateless synthetic tool environments for multi-turn tool interactions (#2410, [#2457])
  • Bring your own inference engine/provider for synthesis (#2504)
  • Reusable ConversationSynthesizer building blocks to drive synthesis programmatically (#2531)
  • Token-usage accounting across a synthesis run (#2480)
  • Misc improvements: [#2559], [#2630], [#2588], [#2591], [#2576]

🔁 Partial-failure-tolerant pipelines

Long inference, synthesis, and judging runs can now return partial results instead of failing the whole batch. New infer_partial() / synthesize_partial() / judge_partial() templates, per-row partial failures in RemoteInferenceEngine, partial-inference result types, and a progress-file reporter you can poll while a run is in flight. (#2498, [#2499], [#2500], [#2501], [#2502])

⚡ Inference hardening

  • Anthropic: configurable prompt-cache TTL, cache tokens folded into prompt_tokens, correct model-version gating for round and dated model names, and sampling params omitted for models that reject them. (#2665, [#2620], [#2568], [#2538])
  • Reliability: honor Retry-After, fixed retry backoff under rate limits, and retries re-paced through the adaptive concurrency controller (which now recovers from backoff and ignores non-retriable errors). (#2516, [#2515], [#2508], [#2507])
  • Token-usage reporting from the vLLM engine, reasoning_content parsing from remote responses, user_id forwarding for abuse attribution, preserved Accept-Encoding across engine header overrides, and a new HuggingFace remote inference engine. (#2606, [#2465], [#2519], [#2622], [#2434])
  • Tool calling wired through Anthropic, Gemini, Vertex, and vLLM remote requests. (#2420, [#2619])

🚀 Launcher: Modal + SkyPilot-routed Slurm

  • Modal GPU launcher provider (cloud, cluster, client, and batched log retrieval). (#2443, [#2455], [#2503])
  • SkyPilot-routed Slurm (sky-slurm), direct squeue/scontrol job status, forwarded accelerators/cpus/memory to sbatch, JobStatus.submit_time, config-override passthrough, a provider-agnostic ClusterNotFoundError, and several Slurm robustness fixes (unreachable controller and command timeout map to ClusterUnreachableError; COMPLETING maps to RUNNING). (#2477, [#2481], [#2482], [#2506], [#2453], [#2442], [#2535], [#2547], [#2662], [#2493])

🧩 Models, recipes & examples

  • Gemma 4 recipes: 12B / 26B (MoE) / 31B and E2B / E4B, in both FFT and LoRA, with regex LoRA targeting and disk-safe saving; plus a verl buffer-sync patch for robust multi-rank Gemma 4 training. (#2490, [#2491], [#2492], [#2495], [#2496], [#2479], [#2647])
  • Qwen3.5 dual-mode checkpoints get a text_only opt-in, with dual-mode/VLM detection fixed under Transformers 5. (#2557, [#2661])
  • New MedQA example (SFT + Qwen3.5-4B GRPO), the verl Countdown GRPO example fixed to be runnable and checkpoint-producing, and a library tool-use synthesis example. (#2648, [#2656], [#2657], [#2651])
  • Default pad token for Nemotron 3 Nano 4B; muse_glimmer added to the internal model type map. (#2621, [#2595])

📦 Dependency push

transformers → <5.17, trl → <1.7, vllm → <0.24, torch → <2.13, deepspeed → <0.20, plus kernels, torchvision, uvicorn, click, aiohttp, and a pillow security bump. lm_eval is pinned to <0.4.12 — 0.4.12+ breaks group-task evals (e.g. MMLU); single-task evals are unaffected. (#2638, [#2637], [#2511], [#2460], [#2445], [#2447], [#2489], [#2472], [#2475], [#2488], [#2548], [#2645])

Other improvements & fixes

  • Training: fractional eval steps for Hugging Face trainers (#2633), pad_to_multiple_of on TextCompletionsCollatorWithPadding (#2540), clamp auto-detected response template to the generation config (#2600), FSDP merge handling of dim-0 tensors (#2618), preserve _name_or_path in saved config.json (#2470), guard torch._dynamo recompile limit for torch 2.6 (#2549), resolve offline model loads at the pinned revision (#2523, [#2524]).
  • Deploy (Fireworks): deploy by validated deployment shape (#2487), list deployment shapes (#2466), forward idle-window kwargs (#2476, [#2530]), map EXPIRED job state (#2471), and quieter logging (#2469, [#2543]).
  • Datasets & evals: fix preprocessing for raw conversational DPO datasets (#2636), adapt lm_harness to the lm_eval 0.4.12 EvalResults TypedDict (#2450).
  • Inference correctness: clear error on 2xx bodies that parse to JSON null (#2593), guard empty choices / message=None in vLLM (#2459), fix parsed content without tool calls (#2617), per-infer() scratch-file isolation (#2632, [#2634]), Triton GDN-prefill fallback on CUDA < 12.6 (#2518).
  • CI & chores: CodeQL workflow (#2641), coverage config + coverage-unit target (#2438), verl tests (#2437), agent instructions with a publication-approval policy (#2584), doc/link/citation fixes and README news updates (#2639, [#2608], [#2514], [#2541], [#2436]).

New Contributors

  • @lucaszhu-hue made their first contribution in [#2427]
  • @qizwiz made their first contribution in [#2459]
  • @nishilfaldu made their first contribution in [#2639]
  • @androna-xm made their first contribution in [#2657]

Full Changelog: https://github.com/oumi-ai/oumi/compare/v0.8...v0.9

Source: README.md, updated 2026-09-16