Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Tiny vision language model
scikit-learn compatible tabular foundation model
Open image model at the forefront of design
A 0.1B Omni model trained from scratch
Hunyuan Translation Model Version 1.5
Video understanding codebase from FAIR for reproducing video models
Audio foundation model excelling in audio understanding
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
DeepSeek Coder: Let the Code Write Itself
Qwen2.5-VL is the multimodal large language model series
1B text generation model based on the HRM architecture
Robust Speech Recognition Across Languages, Dialects
Repo for SeedVR2 & SeedVR
Ling-V2 is a MoE LLM provided and open-sourced by InclusionAI
Open-Source Financial Large Language Models
A Customizable Image-to-Video Model based on HunyuanVideo
Generate Any 3D Scene in Seconds
Repo of Qwen2-Audio chat & pretrained large audio language model
Tongyi Deep Research, the Leading Open-source Deep Research Agent
The Clay Foundation Model - An open source AI model and interface
Advancing Open-source World Models
code for Mesh R-CNN, ICCV 2019
Designed for text embedding and ranking tasks
Generating Immersive, Explorable, and Interactive 3D Worlds