NeMo AutoModel is NVIDIA's open-source PyTorch Distributed training library for scaling LLM, VLM, diffusion, and retrieval-model training. Its DTensor-native SPMD approach lets the same training code scale from one GPU to large multi-node clusters by changing configuration. Hugging Face integration provides broad model compatibility without requiring format conversion. YAML recipes and CLI overrides keep experiments concise while preserving reproducibility. The library supports composable parallelism, optimized kernels, MoE acceleration, mixed precision, sequence packing, and asynchronous checkpointing. Jobs can run through interactive environments, Slurm, SkyPilot, or Kubernetes-based workflows. It targets both rapid research experiments and high-performance large-scale fine-tuning.

Features

  • DTensor-native SPMD distributed training
  • Hugging Face model compatibility
  • YAML recipes with CLI overrides
  • Composable tensor and pipeline parallelism
  • MoE, mixed-precision, and kernel optimizations
  • Slurm, SkyPilot, and Kubernetes execution

Project Samples

Project Activity

See All Activity >

License

Apache License V2.0

Follow NeMo Automodel

NeMo Automodel Web Site

Other Useful Business Software
Ship Agents Faster Icon
Ship Agents Faster

Transform your applications and workflows into powerful agentic systems at global scale.

Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
Get Started Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of NeMo Automodel!

Additional Project Details

Programming Language

Python

Related Categories

Python Artificial Intelligence Software

Registered

3 days ago