tiktoken is a fast BPE tokeniser for use with OpenAI's models
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
A SOTA open-source image editing model
OCR expert VLM powered by Hunyuan's native multimodal architecture
Qwen3-ASR is an open-source series of ASR models
Audio foundation model excelling in audio understanding
CLIP, Predict the most relevant text snippet given an image
Pokee Deep Research Model Open Source Repo
gpt-oss-120b and gpt-oss-20b are two open-weight language models
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Advancing Open-source World Models
A Systematic Framework for Interactive World Modeling
Models for object and human mesh reconstruction
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
High-Resolution Image Synthesis with Latent Diffusion Models
Tongyi Deep Research, the Leading Open-source Deep Research Agent
Open-source image generative foundation model
Open Source Speech Language Model
Open-source industrial-grade ASR models
Ling is a MoE LLM provided and open-sourced by InclusionAI
Multimodal Diffusion with Representation Alignment
General-purpose image editing model that delivers high-fidelity
Open-source deep-learning framework
The Clay Foundation Model - An open source AI model and interface
An Open Real-time Video-Language Interaction System