Audio Language Models are Few-Shot Learners
Flux 2 image generation model pure C inference
Repo for SeedVR2 & SeedVR
FAIR Sequence Modeling Toolkit 2
Advancing Open-source World Models
Controllable & emotion-expressive zero-shot TTS
Hunyuan Translation Model Version 1.5
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Open-source multi-speaker long-form text-to-speech model
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Open-source large language model family from Tencent Hunyuan
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Open-weight, large-scale hybrid-attention reasoning model
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
LLM-based Reinforcement Learning audio edit model
Multimodal embedding and reranking models built on Qwen3-VL
Official implementation of Watermark Anything with Localized Messages
General-purpose image editing model that delivers high-fidelity
Ling-V2 is a MoE LLM provided and open-sourced by InclusionAI
Reproduction of Poetiq's record-breaking submission to the ARC-AGI-1
Language modeling in a sentence representation space
Designed for text embedding and ranking tasks
Multi-modal large language model designed for audio understanding