A Powerful Native Multimodal Model for Image Generation
High-Resolution Image Synthesis with Latent Diffusion Models
A theoretical reconstruction of the Claude Mythos architecture
Fast and Universal 3D reconstruction model for versatile tasks
Diffusion Transformer with Fine-Grained Chinese Understanding
Industrial-level controllable zero-shot text-to-speech system
Hackable and optimized Transformers building blocks
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Contexts Optical Compression
gpt-oss-120b and gpt-oss-20b are two open-weight language models
Fast-stable-diffusion + DreamBooth
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
Multimodal-Driven Architecture for Customized Video Generation
Open Source Speech Language Model
Z80-μLM is a 2-bit quantized language model
DeepSeek Coder: Let the Code Write Itself
A 0.1B Omni model trained from scratch
26m function call model that runs on incredibly small devices
Tongyi Deep Research, the Leading Open-source Deep Research Agent
4M: Massively Multimodal Masked Modeling
Open-source deep-learning framework
Provides convenient access to the Anthropic REST API from any Python 3
Miso TTS is an 8 billion, highly emotive text-to-speech model
Designed for text embedding and ranking tasks
Open image model at the forefront of design