Mixture-of-Experts Vision-Language Models for Advanced Multimodal
Multimodal Diffusion with Representation Alignment
Multimodal-Driven Architecture for Customized Video Generation
Open Source Speech Language Model
A Family of Open Sourced Music Foundation Models
Strong, Economical, and Efficient Mixture-of-Experts Language Model
Agentic, Reasoning, and Coding (ARC) foundation models
CLIP, Predict the most relevant text snippet given an image
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image
RGBD video generation model conditioned on camera input
Analyze computation-communication overlap in V3/R1
Open Multilingual Multimodal Chat LMs
Official code for Style Aligned Image Generation via Shared Attention
A library for Multilingual Unsupervised or Supervised word Embeddings
CTC-based forced aligner for audio-text in 158 languages
Qwen3-Next: 80B instruct LLM with ultra-long context up to 1M tokens
Hermes 4 FP8: hybrid reasoning Llama-3.1-405B model by Nous Research
Multimodal 7B model for image, video, and text understanding tasks
Instruction-tuned 1.2B LLM for multilingual text generation by Meta
CLIP model fine-tuned for zero-shot fashion product classification