4M is a training framework for “any-to-any” vision foundation models that uses tokenization and masking to scale across many modalities and tasks. The same model family can classify, segment, detect, caption, and even generate images, with a single interface for both discriminative and generative use. The repository releases code and models for multiple variants (e.g., 4M-7 and 4M-21), emphasizing transfer to unseen tasks and modalities. Training/inference configs and issues discuss things like depth tokenizers, input masks for generation, and CUDA build questions, signaling active research iteration. The design leans into flexibility and steerability, so prompts and masks can shape behavior without bespoke heads per task. In short, 4M provides a unified recipe to pretrain large multimodal models that generalize broadly while remaining practical to fine-tune.

Features

  • Any-to-any modeling across diverse vision tasks
  • Masked modeling with unified tokenization for multiple modalities
  • Released model families (e.g., 4M-7, 4M-21) with training/eval code
  • Promptable and steerable behavior without task-specific heads
  • Transfer to unseen tasks and modalities from a single backbone
  • Research-grade configs and examples for reproduction

Project Samples

Project Activity

See All Activity >

Categories

AI Models

License

Apache License V2.0

Follow 4M

4M Web Site

Other Useful Business Software
Veeam Data Platform v13.1 - Get Your Free Trial Icon
Veeam Data Platform v13.1 - Get Your Free Trial

Secure by design, portable by default. Recover clean, fast, anywhere. Start a free trial.

Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
Try it Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of 4M!

Additional Project Details

Programming Language

Python

Related Categories

Python AI Models

Registered

2025-10-08