Phi-3.5 for Mac: Locally-run Vision and Language Models
State-of-the-art TTS model under 25MB
Large Multimodal Models for Video Understanding and Editing
Code for running inference and finetuning with SAM 3 model
Contexts Optical Compression
A Family of Open Sourced Music Foundation Models
Agentic, Reasoning, and Coding (ARC) foundation models
AlphaFold 3 inference pipeline
Open Source Speech Language Model
Multimodal-Driven Architecture for Customized Video Generation
Multimodal Diffusion with Representation Alignment
Visual Causal Flow
Generate Any 3D Scene in Seconds
A Production-ready Reinforcement Learning AI Agent Library
Convert Google Gemini web into OpenAI-compatible API
A theoretical reconstruction of the Claude Mythos architecture
Research code artifacts for Code World Model (CWM)
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Qwen3-ASR is an open-source series of ASR models
Qwen3 is the large language model series developed by Qwen team
Unified Multimodal Understanding and Generation Models
VMZ: Model Zoo for Video Modeling
Implementation of the Surya Foundation Model for Heliophysics
Netease Youdao's open-source embedding and reranker models
Inference code for scalable emulation of protein equilibrium ensembles