MiniMax H3 is a general-purpose, omni-modal generative system
A Multi-Modal World Model for Reconstructing, Generating, Simulation
Powerful AI language model (MoE) optimized for efficiency/performance
Long-form streaming TTS system for multi-speaker dialogue generation
Industrial-level controllable zero-shot text-to-speech system
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Open-source multi-speaker long-form text-to-speech model
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
Official repository for LTX-Video
Robust Speech Recognition Across Languages, Dialects
Official inference repo for FLUX.2 models
Controllable & emotion-expressive zero-shot TTS
Qwen-Image is a powerful image generation foundation model
Wan2.2: Open and Advanced Large-Scale Video Generative Model
General-purpose image editing model that delivers high-fidelity
Models for object and human mesh reconstruction
Video Object and Interaction Deletion
Diffusion Transformer with Fine-Grained Chinese Understanding
Pokee Deep Research Model Open Source Repo
FAIR Sequence Modeling Toolkit 2
Qwen2.5-VL is the multimodal large language model series
New family of code large language models (LLMs)
AI cognitive-enhancement Skills based on Anthropic's J-space
High-Resolution 3D Assets Generation with Large Scale Diffusion Models
Reference PyTorch implementation and models for DINOv3