Download Latest Version v3.0.0 source code.zip (2.2 MB) Google Add to Preferred Sources
Home / v2.4.0
Name Modified Size InfoDownloads / Week
Parent folder
ai2_olmo_core-2.4.0-py3-none-any.whl 2025-11-20 582.4 kB
ai2_olmo_core-2.4.0.tar.gz 2025-11-20 496.6 kB
README.md 2025-11-20 3.6 kB
v2.4.0 source code.tar.gz 2025-11-20 757.4 kB
v2.4.0 source code.zip 2025-11-20 935.3 kB
Totals: 5 Items   2.8 MB 0

What's new

Added 🎉

  • Added option to skip ranges of steps in the trainer.
  • Send a Slack notification when a Beaker job appears to be stuck.
  • Added ignore_fingerprint_mismatch parameter to NumpyDataLoaderConfig to allow resuming training from a checkpoint with a different dataset mix.
  • Added helpful error messages when OLMo-mix-0625 files are not found, directing users to use OLMo-mix-0925 and the fingerprint override flag.
  • Added olmo_core.generate.chat module to allow interacting with OlmoCore models without conversion to other formats.
  • Added GAPMonitorCallback for monitoring gradients, activations, and parameters (GAP).
  • Added official Olmo 3 7B and 32B pretraining scripts and data mix.
  • Added official Olmo 3 7B and 32B midtraining scripts and data mix.
  • Added official Olmo 3 7B and 32B long-context scripts and data mix.
  • Added a NoOpOptimizer that does nothing, uses no memory, and can be used for debugging.
  • Added official config for Olmo 3 32B.
  • Olmo 3 model card and checkpoint manifests.

Fixed ✅

  • Set missing NCCL_NVLSTREE_MAX_CHUNKSIZE env var that is now needed for running jobs on Augusta cluster.
  • Fixed bug with RemoteFileSystemReader that caused excess memory usage.
  • No longer overrides random's RNG seed when building SourceMixtureDatasetConfig.
  • Fix handling URLs in olmo_core.nn.hf.checkpoint.save_hf_model and in examples/huggingface.
  • Fix potential NaN loss that can occur when using instance masking.
  • Stability improvements developed while training Olmo3 32B.

Changed ⚠️

  • Removed unused field in YaRNRoPEScalingConfig.

Commits

1ed8900e (chore) prepare for release v2.4.0 2c179c2b (chore) prepare for release v2.4.0 (#467) 7e0431f9 Fix link to 7B midtrain script (#469) 843fe3d3 Olmo3 model cards, checkpoint manifest, and readme (#468) cbdc2f1a Olmo3 32B cleanup and checkin (#460) 20548a05 Official Olmo3 32B long-context script (#465) 14b15cc2 Official Olmo3 32B midtrain script(s) and mix(es) (#466) a25a5144 Official Olmo3 32B pretrain config and data mix (#464) 6b73ba05 Official Olmo3-7B long context script (#458) bdc61e4a Official Olmo3-7B midtraining scripts (#445) 55804bf8 32B official config (#454) a86131dd Slight refactor of Yarn Scaling Config (#456) 68c74093 Handle target URLs properly in HF conversion (#453) 2504cc28 Instance mask correction to avoid nan loss (#452) 137274ec Add callback to monitor grads, activations, params (#446) 0959a548 Improve mem usage of RemoteFileSystemReader (#451) 600d2fe2 Official Olmo3-7B pretraining scripts (#443) accc3100 make launch timeout configurable from CLI aa0e6290 Avoid overriding RNG seed when building SourceMixtureDatasetConfig (#449) 98ba2e41 NoOp optimizer (#444) 03e68365 OlmoCore native chat interface (#439) 7a0bbd76 unset 2 NCCL env vars per Google's recommendation aacb6eba only send local Slack notifications when callback is enabled (#441) bfc8d7ab Min python version to 3.10 (#442) 96d43d4e Set missing NCCL_NVLSTREE_MAX_CHUNKSIZE env var (#440) 5ad6db5e hot fix for listing gcs dirs 21869570 Allow manual bypass of fingerprint mismatch when switching datasets (#435) 043505da hot fix to step regex dd7e747c Send a Slack notification when a Beaker job appears to be stuck (#431) e27a9b40 Add WSDS (Warmup-Stable-Decay-Simplified) Scheduler (#419) c92320f2 Use a dataclass for 'Trainer.steps_to_skip' (#430) 9669268a clean up checkpointing code to minimize distributed communication (#428) 87d64b95 fix changelog 269bf022 Add option to skip ranges of steps in the trainer (#425)

Source: README.md, updated 2025-11-20