| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| README.md | 2026-09-16 | 8.4 kB | |
| v0.9 source code.tar.gz | 2026-09-16 | 3.5 MB | |
| v0.9 source code.zip | 2026-09-16 | 4.7 MB | |
| Totals: 3 Items | 8.3 MB | 0 | |
Oumi v0.9 Release Notes
This release makes Oumi an end-to-end stack for agentic, tool-using models: you can now train on tool-calling data, execute tool calls against real environments (databases, HTTP endpoints, deterministic lookups, and model-simulated tools), synthesize multi-turn tool conversations, and run GRPO/RL over any of those environments with verl. It also adds a multi-criteria RubricJudge, partial-failure-tolerant inference/synthesis/judging, a Modal GPU launcher and SkyPilot-routed Slurm, and new Gemma 4 recipes.
Highlights
🛠️ Tool-use training
Train models to call tools end-to-end:
- Tool-aware chat templates — access the tool-call section, auto-detect the tool-call form a template expects (#2605, [#2603], [#2604])
- Trainable tool adapters (#2658)
- Collator masking that correctly remasks tool-response spans, including responses nested in an assistant turn (#2594)
- Tool-calling SFT and text DPO over JSONL, with serialized tool arguments decoded during tokenization (#2642, [#2517])
- Canonical tool-call
Conversationshape + JSON-string ↔ dict argument helpers (#2602, [#2601])
🌐 Tool-execution Environments
The new oumi.environments subsystem supports executing tools calls with various backends.
ExecutableEnvironment/ExecutableTool— the base abstraction for tools that actually run. (#2527)DatabaseExecutableEnvironment— run tool calls against a database session, with rollback-isolated teardown (see the EHR example). (#2528, [#2545])EndpointEnvironment— dispatch tool calls to HTTP endpoints. (#2625)- Lookup (deterministic) and simulated environments — deterministic lookup tables validated against tool schemas, with argument defaults applied and any JSON value as output; model-simulated tool execution for synthesis. (#2623, [#2536], [#2573], [#2574])
- Tool executors are resolvable by registry name, so environments can wire tools declaratively. (#2555)
🎯 OumiVerlTool — RL over any Oumi environment
OumiVerlTool bridges Oumi's executable environments into verl, so you can run verl GRPO tool-use RL over any Oumi environment. Custom judge rewards are now supported in verl GRPO, multi-turn prompts are supported in the GRPO data path, conversation metadata is preserved through GRPO datasets, and per-run verl metrics are mirrored to verl_metrics.jsonl with richer logging.
(#2569, [#2624], [#2542], [#2644], [#2643], [#2646])
⚖️ RubricJudge — multi-criteria evaluation
A new RubricJudge scores a response against several named criteria in a single judge call, instead of one boolean per pass — useful both for evaluation and as a GRPO reward. Ships with example configs under configs/projects/judges/rubric and documentation.
(#2626, [#2628], [#2629])
🧬 Conversation synthesis
Synthesize multi-turn tool-use conversations:
- Stateful and stateless synthetic tool environments for multi-turn tool interactions (#2410, [#2457])
- Bring your own inference engine/provider for synthesis (#2504)
- Reusable
ConversationSynthesizerbuilding blocks to drive synthesis programmatically (#2531) - Token-usage accounting across a synthesis run (#2480)
- Misc improvements: [#2559], [#2630], [#2588], [#2591], [#2576]
🔁 Partial-failure-tolerant pipelines
Long inference, synthesis, and judging runs can now return partial results instead of failing the whole batch. New infer_partial() / synthesize_partial() / judge_partial() templates, per-row partial failures in RemoteInferenceEngine, partial-inference result types, and a progress-file reporter you can poll while a run is in flight.
(#2498, [#2499], [#2500], [#2501], [#2502])
⚡ Inference hardening
- Anthropic: configurable prompt-cache TTL, cache tokens folded into
prompt_tokens, correct model-version gating for round and dated model names, and sampling params omitted for models that reject them. (#2665, [#2620], [#2568], [#2538]) - Reliability: honor
Retry-After, fixed retry backoff under rate limits, and retries re-paced through the adaptive concurrency controller (which now recovers from backoff and ignores non-retriable errors). (#2516, [#2515], [#2508], [#2507]) - Token-usage reporting from the vLLM engine,
reasoning_contentparsing from remote responses,user_idforwarding for abuse attribution, preservedAccept-Encodingacross engine header overrides, and a new HuggingFace remote inference engine. (#2606, [#2465], [#2519], [#2622], [#2434]) - Tool calling wired through Anthropic, Gemini, Vertex, and vLLM remote requests. (#2420, [#2619])
🚀 Launcher: Modal + SkyPilot-routed Slurm
- Modal GPU launcher provider (cloud, cluster, client, and batched log retrieval). (#2443, [#2455], [#2503])
- SkyPilot-routed Slurm (
sky-slurm), directsqueue/scontroljob status, forwardedaccelerators/cpus/memorytosbatch,JobStatus.submit_time, config-override passthrough, a provider-agnosticClusterNotFoundError, and several Slurm robustness fixes (unreachable controller and command timeout map toClusterUnreachableError;COMPLETINGmaps toRUNNING). (#2477, [#2481], [#2482], [#2506], [#2453], [#2442], [#2535], [#2547], [#2662], [#2493])
🧩 Models, recipes & examples
- Gemma 4 recipes: 12B / 26B (MoE) / 31B and E2B / E4B, in both FFT and LoRA, with regex LoRA targeting and disk-safe saving; plus a verl buffer-sync patch for robust multi-rank Gemma 4 training. (#2490, [#2491], [#2492], [#2495], [#2496], [#2479], [#2647])
- Qwen3.5 dual-mode checkpoints get a
text_onlyopt-in, with dual-mode/VLM detection fixed under Transformers 5. (#2557, [#2661]) - New MedQA example (SFT + Qwen3.5-4B GRPO), the verl Countdown GRPO example fixed to be runnable and checkpoint-producing, and a library tool-use synthesis example. (#2648, [#2656], [#2657], [#2651])
- Default pad token for Nemotron 3 Nano 4B;
muse_glimmeradded to the internal model type map. (#2621, [#2595])
📦 Dependency push
transformers → <5.17, trl → <1.7, vllm → <0.24, torch → <2.13, deepspeed → <0.20, plus kernels, torchvision, uvicorn, click, aiohttp, and a pillow security bump. lm_eval is pinned to <0.4.12 — 0.4.12+ breaks group-task evals (e.g. MMLU); single-task evals are unaffected.
(#2638, [#2637], [#2511], [#2460], [#2445], [#2447], [#2489], [#2472], [#2475], [#2488], [#2548], [#2645])
Other improvements & fixes
- Training: fractional eval steps for Hugging Face trainers (#2633),
pad_to_multiple_ofonTextCompletionsCollatorWithPadding(#2540), clamp auto-detected response template to the generation config (#2600), FSDP merge handling of dim-0 tensors (#2618), preserve_name_or_pathin savedconfig.json(#2470), guardtorch._dynamorecompile limit for torch 2.6 (#2549), resolve offline model loads at the pinned revision (#2523, [#2524]). - Deploy (Fireworks): deploy by validated deployment shape (#2487), list deployment shapes (#2466), forward idle-window kwargs (#2476, [#2530]), map EXPIRED job state (#2471), and quieter logging (#2469, [#2543]).
- Datasets & evals: fix preprocessing for raw conversational DPO datasets (#2636), adapt
lm_harnessto the lm_eval 0.4.12EvalResultsTypedDict (#2450). - Inference correctness: clear error on 2xx bodies that parse to JSON null (#2593), guard empty choices /
message=Nonein vLLM (#2459), fix parsed content without tool calls (#2617), per-infer()scratch-file isolation (#2632, [#2634]), Triton GDN-prefill fallback on CUDA < 12.6 (#2518). - CI & chores: CodeQL workflow (#2641), coverage config +
coverage-unittarget (#2438), verl tests (#2437), agent instructions with a publication-approval policy (#2584), doc/link/citation fixes and README news updates (#2639, [#2608], [#2514], [#2541], [#2436]).
New Contributors
- @lucaszhu-hue made their first contribution in [#2427]
- @qizwiz made their first contribution in [#2459]
- @nishilfaldu made their first contribution in [#2639]
- @androna-xm made their first contribution in [#2657]
Full Changelog: https://github.com/oumi-ai/oumi/compare/v0.8...v0.9