Marin is an open-source research platform and community for developing foundation models through transparent, reproducible experimentation. It covers the complete model-building pipeline from data curation and filtering through tokenization, pretraining, post-training, and evaluation. Experiments and decisions are documented as they occur, including unsuccessful approaches. The framework is primarily used for large language models but has also supported audio-text, DNA, and protein modeling research. Experiments are expressed as dependent steps that execute in topological order, enabling reproducible training workflows. Marin also provides model checkpoints, training recipes, scaling research, documentation, and reusable infrastructure for large-scale experiments.

Features

  • Foundation model training workflows
  • Data curation and tokenization pipelines
  • Pretraining and post-training support
  • Model evaluation infrastructure
  • Dependency-based experiment execution
  • Open checkpoints and reproducible research artifacts

Project Samples

Project Activity

See All Activity >

Categories

Frameworks

License

Apache License V2.0

Follow Marin

Marin Web Site

Other Useful Business Software
Veeam Data Platform v13.1 Icon
Veeam Data Platform v13.1

Move workloads across hypervisors and clouds with no vendor lock-in. Try VDP free today.

Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
Try Now
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of Marin!

Additional Project Details

Programming Language

Python

Related Categories

Python Frameworks

Registered

2026-08-27