Open-weight, large-scale hybrid-attention reasoning model
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Long-form streaming TTS system for multi-speaker dialogue generation
New family of code large language models (LLMs)
An AI-powered security review GitHub Action using Claude
GLM-4-Voice | End-to-End Chinese-English Conversational Model
Renderer for the harmony response format to be used with gpt-oss
1B text generation model based on the HRM architecture
Open image model at the forefront of design
Hunyuan Translation Model Version 1.5
Capable of understanding text, audio, vision, video
General-purpose image editing model that delivers high-fidelity
Open-source multi-speaker long-form text-to-speech model
Qwen3-omni is a natively end-to-end, omni-modal LLM
Ling-V2 is a MoE LLM provided and open-sourced by InclusionAI
Generate Any 3D Scene in Seconds
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Unified Multimodal Understanding and Generation Models
This repository contains the official implementation of FastVLM
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training
Block Diffusion for Ultra-Fast Speculative Decoding
High-resolution models for human tasks
ICLR2024 Spotlight: curation/training code, metadata, distribution
Diffusion Transformer with Fine-Grained Chinese Understanding
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming