Sharp Monocular Metric Depth in Less Than a Second
A theoretical reconstruction of the Claude Mythos architecture
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Industrial-level controllable zero-shot text-to-speech system
Unified Multimodal Understanding and Generation Models
A Powerful Native Multimodal Model for Image Generation
Open-source multi-speaker long-form text-to-speech model
Open-source image generative foundation model
OpenTinker is an RL-as-a-Service infrastructure for foundation models
An experimental version of DeepSeek model
Infinite Worlds with Versatile Interactions
A 0.1B Omni model trained from scratch
Fast-stable-diffusion + DreamBooth
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
DeepSeek Coder: Let the Code Write Itself
Open-source deep-learning framework
Foundation Models for Time Series
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Generate Any 3D Scene in Seconds
Audio foundation model excelling in audio understanding
A Systematic Framework for Interactive World Modeling
Open-Source Financial Large Language Models
Inference script for Oasis 500M
Video Object and Interaction Deletion
High-resolution models for human tasks