A theoretical reconstruction of the Claude Mythos architecture
Qwen-Image is a powerful image generation foundation model
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Z80-μLM is a 2-bit quantized language model
Open-source image generative foundation model
Sharp Monocular Metric Depth in Less Than a Second
High-Resolution Image Synthesis with Latent Diffusion Models
An experimental version of DeepSeek model
Open-source multi-speaker long-form text-to-speech model
Infinite Worlds with Versatile Interactions
General-purpose image editing model that delivers high-fidelity
Accurate × Fast × Comprehensive
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Unified Multimodal Understanding and Generation Models
GLM-4 series: Open Multilingual Multimodal Chat LMs
Phi-3.5 for Mac: Locally-run Vision and Language Models
Long-form streaming TTS system for multi-speaker dialogue generation
Inference script for Oasis 500M
A Systematic Framework for Interactive World Modeling
Industrial-level controllable zero-shot text-to-speech system
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Fast-stable-diffusion + DreamBooth
Multimodal-Driven Architecture for Customized Video Generation
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence