| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| ai2_olmo_core-2.4.0-py3-none-any.whl | 2025-11-20 | 582.4 kB | |
| ai2_olmo_core-2.4.0.tar.gz | 2025-11-20 | 496.6 kB | |
| README.md | 2025-11-20 | 3.6 kB | |
| v2.4.0 source code.tar.gz | 2025-11-20 | 757.4 kB | |
| v2.4.0 source code.zip | 2025-11-20 | 935.3 kB | |
| Totals: 5 Items | 2.8 MB | 0 | |
What's new
Added 🎉
- Added option to skip ranges of steps in the trainer.
- Send a Slack notification when a Beaker job appears to be stuck.
- Added
ignore_fingerprint_mismatchparameter toNumpyDataLoaderConfigto allow resuming training from a checkpoint with a different dataset mix. - Added helpful error messages when OLMo-mix-0625 files are not found, directing users to use OLMo-mix-0925 and the fingerprint override flag.
- Added
olmo_core.generate.chatmodule to allow interacting with OlmoCore models without conversion to other formats. - Added
GAPMonitorCallbackfor monitoring gradients, activations, and parameters (GAP). - Added official Olmo 3 7B and 32B pretraining scripts and data mix.
- Added official Olmo 3 7B and 32B midtraining scripts and data mix.
- Added official Olmo 3 7B and 32B long-context scripts and data mix.
- Added a
NoOpOptimizerthat does nothing, uses no memory, and can be used for debugging. - Added official config for Olmo 3 32B.
- Olmo 3 model card and checkpoint manifests.
Fixed ✅
- Set missing
NCCL_NVLSTREE_MAX_CHUNKSIZEenv var that is now needed for running jobs on Augusta cluster. - Fixed bug with
RemoteFileSystemReaderthat caused excess memory usage. - No longer overrides
random's RNG seed when buildingSourceMixtureDatasetConfig. - Fix handling URLs in
olmo_core.nn.hf.checkpoint.save_hf_modeland inexamples/huggingface. - Fix potential NaN loss that can occur when using instance masking.
- Stability improvements developed while training Olmo3 32B.
Changed ⚠️
- Removed unused field in
YaRNRoPEScalingConfig.
Commits
1ed8900e (chore) prepare for release v2.4.0
2c179c2b (chore) prepare for release v2.4.0 (#467)
7e0431f9 Fix link to 7B midtrain script (#469)
843fe3d3 Olmo3 model cards, checkpoint manifest, and readme (#468)
cbdc2f1a Olmo3 32B cleanup and checkin (#460)
20548a05 Official Olmo3 32B long-context script (#465)
14b15cc2 Official Olmo3 32B midtrain script(s) and mix(es) (#466)
a25a5144 Official Olmo3 32B pretrain config and data mix (#464)
6b73ba05 Official Olmo3-7B long context script (#458)
bdc61e4a Official Olmo3-7B midtraining scripts (#445)
55804bf8 32B official config (#454)
a86131dd Slight refactor of Yarn Scaling Config (#456)
68c74093 Handle target URLs properly in HF conversion (#453)
2504cc28 Instance mask correction to avoid nan loss (#452)
137274ec Add callback to monitor grads, activations, params (#446)
0959a548 Improve mem usage of RemoteFileSystemReader (#451)
600d2fe2 Official Olmo3-7B pretraining scripts (#443)
accc3100 make launch timeout configurable from CLI
aa0e6290 Avoid overriding RNG seed when building SourceMixtureDatasetConfig (#449)
98ba2e41 NoOp optimizer (#444)
03e68365 OlmoCore native chat interface (#439)
7a0bbd76 unset 2 NCCL env vars per Google's recommendation
aacb6eba only send local Slack notifications when callback is enabled (#441)
bfc8d7ab Min python version to 3.10 (#442)
96d43d4e Set missing NCCL_NVLSTREE_MAX_CHUNKSIZE env var (#440)
5ad6db5e hot fix for listing gcs dirs
21869570 Allow manual bypass of fingerprint mismatch when switching datasets (#435)
043505da hot fix to step regex
dd7e747c Send a Slack notification when a Beaker job appears to be stuck (#431)
e27a9b40 Add WSDS (Warmup-Stable-Decay-Simplified) Scheduler (#419)
c92320f2 Use a dataclass for 'Trainer.steps_to_skip' (#430)
9669268a clean up checkpointing code to minimize distributed communication (#428)
87d64b95 fix changelog
269bf022 Add option to skip ranges of steps in the trainer (#425)