| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| ai2_olmo_core-2.6.0-py3-none-any.whl | 2026-08-11 | 750.4 kB | |
| ai2_olmo_core-2.6.0.tar.gz | 2026-08-11 | 632.3 kB | |
| README.md | 2026-08-11 | 8.1 kB | |
| v2.6.0 source code.tar.gz | 2026-08-11 | 1.1 MB | |
| v2.6.0 source code.zip | 2026-08-11 | 1.3 MB | |
| Totals: 5 Items | 3.8 MB | 0 | |
What's new
Added 🎉
- Added
max_checkpointsparameter toCheckpointerCallback(default: 3) to limit the number of permanent checkpoints retained. Oldest checkpoints are removed automatically when the limit is exceeded. Set toNoneto keep all (previous behavior). - Added
OutputDiscardCheckpoint, an activation-recompute primitive for cases where the output of a checkpointed region dominates memory rather than its intermediates (e.g. precision casts, FFN up-projections). Forward runs underno_grad, the output's storage can be freed after downstream consumption, and a backward hook recomputes and rebinds the freed storage in place via a C++share_storageextension (with a Python fallback for environments without a C++ toolchain). - Added Qwen3.5 dense model configs (0.8B, 4B, 9B, 27B) with hybrid Gated DeltaNet + full-attention architecture.
- Added partial RoPE support via
partial_rotary_factoron :class:~olmo_core.nn.rope.RoPEConfig. - Added HuggingFace weight conversion for
qwen3_5_texthybrid models. - Added a configurable vision transformer encoder (
VisionTransformer, configured viaVisionEncoderConfig), vision-to-LM connector (VisionConnector), andMultimodalLM— a composite vision-language model that fuses image patch tokens into the LM token stream. Supports OpenAI CLIP, SigLIP, and SigLIP2 encoder variants with factory configs for all standard Molmo2 checkpoints. - Added
HFConverterCallback, which can be used to convert models to huggingface format at the end of the training run. - Trainer now records checkpoint save and load durations as
train/checkpoint_save_duration_sandtrain/checkpoint_load_duration_smetrics. - Added
PowerLR, a power-law learning rate scheduler with linear warmup, power-decay phase (lr = initial_lr * (current / warmup) ** bfor negativeb, making the LR independent of the training horizon), and an optional linear decay tail. Registered as"power_lr". - Added
ComposableScheduler, a piecewise LR scheduler built fromComposableSchedulerStagesegments (linear/cosine interpolation between endpoint LRs) on an absolute time axis. Registered as"composable". Note:ComposableSchedulerignores thet_maxpassed toget_lrand emits a once-per-instanceUserWarningto that effect. - Added
OverrideDecay, a late-stage decay override usable on bothComposableSchedulerandSequentialSchedulervia anoverride_decayfield. Whencurrent >= override_decay.start, the main schedule is interrupted mid-flight and the LR decays from the value the main schedule would have produced atstartto a target LR overduration(linear or cosine).SequentialScheduleradditionally warns thatt_maxis ignored once the override becomes active. OLMO_RICH_LOGGINGcan now explicitly enable or disable rich console logging (0/false/no/offdisables it); previously setting it to any value only force-enabled rich logging.init_distributed()now bootstraps a minimal single-process environment (RANK=0,WORLD_SIZE=1,MASTER_ADDR/MASTER_PORT) when launch env vars are absent, so scripts can be run directly (withouttorchrun) for single-process debugging.- Added a configurable
determinism_checkoption to activation checkpointing (default"default"); set it to"none"to skip torch's recompute metadata check for opaque linear-attention kernels undertorch.compile.
Fixed ✅
- The CPU
TestCI job now cachesHF_HOMEacross runs so the HuggingFace roundtrip tests (Qwen3-0.6B, Gemma-3-270m) don't re-download their checkpoints every run. - Excluded
mark_dynamicfromtorch.compiletracing (@torch.compiler.disable). - Clearer error messages (now include the offending values) when a rank batch size isn't divisible by the sequence length, or
max_target_sequence_lengthisn't a multiple ofsequence_length. - S3 uploads/downloads now also retry on transient SSL errors (
ssl.SSLError, botocore/urllib3SSLError). - Distributed checkpoint writes now clone each tensor before serialization to avoid accidentally writing the full backing storage of a view/shared tensor, with a guard that raises
OLMoCheckpointErrorif a written tensor is unexpectedly larger than itsnbytes. - Fixed LM in-loop evaluator data-order drift across repeated runs by resetting loader bookkeeping before each pass and making deterministic reshuffling the default.
- Fixed Qwen3 implementation to match HuggingFace by applying RoPE in the input dtype (bf16) rather than upcasting to fp32.
- Fixed HF model conversion for Llama, Qwen3, and Gemma so that converted checkpoints roundtrip correctly.
- Fixed Beaker secret existence check to use the case-insensitive HTTP endpoint, avoiding spurious "secret not found" errors when secret names differ only in case.
- Fixed
Transformer.init_weightsso that under interleaved pipeline parallelism (e.g.Interleaved1F1B,InterleavedZeroBubble) the multiple model chunks owned by a single rank no longer initialize to identical parameters. Adds amodel_part_idxkwarg incorporated into the seed asmodel_part_idx * pp_size. - Disabled
torch.compiletracing throughTEAttentionBackend.forward, whose Python/pybind setup is not Dynamo-safe. - Fixed
TransformerPipelineTrainModule.num_flops_per_tokenreturningNoneunder pipeline parallelism. Each PP rank only holds its stage's layers, so summing FLOPs frommodel_partsundercounts the model. Capturemodel.num_flops_per_tokenas a bound method beforesplit_modeldeepcopies and drops layers, then call it at metric time. On meta device (the standard PP init path) this has no memory cost.
Changed ⚠️
- Set transformers version to >= 5.4.0 for Qwen 3.5 and in sync with open-instruct
- Added a documented
deterministicoption toLMEvaluatorandLMEvaluatorCallbackConfigso callers can opt out of fixed eval ordering when desired.
Commits
b7e9671d7 (chore) prepare for release v2.6.0 77714b793 fix dion pypi (#816) 5b0d1cffa (chore) prepare for release v2.6.0 66f768b22 Update anonymous paths in public scripts to point to data in hugging face bucket (#802) 064b172e5 bump fla to 0.5.2 (#798) d3146cc5e Misc fixes (#800) fa6c5014c Add max_checkpoints to limit permanent checkpoint retention (#694) c3802ed54 CI: cache HuggingFace models for the CPU Test job (#732) 2000b1b9b Add OutputDiscardCheckpoint (#682) 9aa3280fd Add Qwen3.5 model support (0.8B, 4B, 9B, 27B) (#684) 36f99f06d Add conversion overrides for Llama, Qwen3, and Gemma 4 models so they roundtrip properly (#677) 885219bfa Add configurable determinism_check to activation checkpointing (#713) 8d22ca94f [1/n] Add vision transformer, connector, and MultimodalLM (#692) 59a339f53 Misc training-utility improvements (#691) 1713ea313 Improve checkpoint/S3 IO robustness (#690) 754d58d8c CI: run GPU tests at low priority (#696) 1af17a441 Fix PP FLOPs: capture full model before pipeline split (#680) 525cc2540 Increase tolerance for flaky test (#681) 38704d167 Two small transformer core fixes: TE Dynamo + PP init seed (#679) 73637f7b8 Add PowerLR scheduler (#674) 2caaee970 Add ComposableScheduler (#671) 2e67bcb08 Update task timeout for GPU tests (#675) e556a86fe Revert "Add position_ids-based varlen RoPE support for packed inputs" (#672) 1eec69671 Add position_ids-based varlen RoPE support for packed inputs (#654) 60d2487ab Fixes the Qwen3 implementation to match HF (#663) 5e7ee43cf Record checkpoint save/load durations as trainer metrics (#665) 362231828 Fix broken path in 32B LC (#670) 3e19fa23f Fix breaking paths (#669) 53c51c561 Use case-insensitive HTTP endpoint to check beaker secret existence (#666) afe99b60a Pin flash-linear-attention version to 0.4.1 (#664) 2e570869b Add HFConverterCallback for end-of-training HuggingFace conversion (#660) beca1f175 Correct dolmino mix urls (#630) b3760776e Fall back to anonymous GCS client when no credentials are available (#659) 60930ef14 Add gradient dumping to GAPMonitorCallback (#438) befb60b83 make lm evaluator data deterministic (#652)