Tool for exploring and debugging transformer model behaviors
CLIP, Predict the most relevant text snippet given an image
OpenTinker is an RL-as-a-Service infrastructure for foundation models
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Tiny vision language model
A theoretical reconstruction of the Claude Mythos architecture
Open-weight, large-scale hybrid-attention reasoning model
Long-form streaming TTS system for multi-speaker dialogue generation
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
1B text generation model based on the HRM architecture
New family of code large language models (LLMs)
An AI-powered security review GitHub Action using Claude
GLM-4-Voice | End-to-End Chinese-English Conversational Model
Renderer for the harmony response format to be used with gpt-oss
Open image model at the forefront of design
Hunyuan Translation Model Version 1.5
Open-source multi-speaker long-form text-to-speech model
Capable of understanding text, audio, vision, video
General-purpose image editing model that delivers high-fidelity
Qwen3-omni is a natively end-to-end, omni-modal LLM
Ling-V2 is a MoE LLM provided and open-sourced by InclusionAI
Generate Any 3D Scene in Seconds
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Multi-modal large language model designed for audio understanding