MiniMax H3 is a general-purpose, omni-modal generative system
Awesome multilingual OCR toolkits based on PaddlePaddle
Industrial-level controllable zero-shot text-to-speech system
From Images to High-Fidelity 3D Assets
Native and Compact Structured Latents for 3D Generation
AlphaFold 3 inference pipeline
An Open Real-time Video-Language Interaction System
Controllable & emotion-expressive zero-shot TTS
A Multi-Modal World Model for Reconstructing, Generating, Simulation
A theoretical reconstruction of the Claude Mythos architecture
Visual Causal Flow
Contexts Optical Compression
Video Object and Interaction Deletion
Robust Speech Recognition Across Languages, Dialects
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Open Source Speech Language Model
Audio foundation model excelling in audio understanding
RGBD video generation model conditioned on camera input
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Bidirectional token-classification model for identifiable info
Long-form streaming TTS system for multi-speaker dialogue generation
AI cognitive-enhancement Skills based on Anthropic's J-space
Qwen3-ASR is an open-source series of ASR models
Reproduction of Poetiq's record-breaking submission to the ARC-AGI-1
General-purpose image editing model that delivers high-fidelity