GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
Open-source multi-speaker long-form text-to-speech model
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Open-source large language model family from Tencent Hunyuan
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Open-weight, large-scale hybrid-attention reasoning model
Multimodal embedding and reranking models built on Qwen3-VL
Official implementation of Watermark Anything with Localized Messages
General-purpose image editing model that delivers high-fidelity
Ling-V2 is a MoE LLM provided and open-sourced by InclusionAI
Reproduction of Poetiq's record-breaking submission to the ARC-AGI-1
Language modeling in a sentence representation space
Tooling for the Common Objects In 3D dataset
Open-source, high-performance Mixture-of-Experts large language model
Open Multilingual Multimodal Chat LMs
Detect faces in an image
Python example app from the OpenAI API quickstart tutorial
Release for Improved Denoising Diffusion Probabilistic Models
Official DeiT repository
Dataset of GPT-2 outputs for research in detection, biases, and more
Code for the paper Hybrid Spectrogram and Waveform Source Separation
PyTorch implementation of VALL-E (Zero-Shot Text-To-Speech)
Implementation of model parallel autoregressive transformers on GPUs