Semi-Structured Agentic Framework. Workflows build themselves
Motion-controllable Video Generation via Latent Trajectory Guidance
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
1B text generation model based on the HRM architecture
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Foundational video generation model with 13.6B parameters
Medical imaging toolkit for deep learning
UI-TARS-desktop version that can operate on your local personal device
Open Source Differentiable Computer Vision Library
SOTA discrete acoustic codec models with 40/75 tokens per second
A tool to use the Ai2 Open Coding Agents Soft-Verified Agents
Multimodal embedding and reranking models built on Qwen3-VL
A fast library for AutoML and tuning
One-ink editorial print image skill
AI-Driven Exploration in the Space of Code
Diffusion Transformer with Fine-Grained Chinese Understanding
High-Fidelity and Controllable Generation of Textured 3D Assets
Large Multimodal Models for Video Understanding and Editing
Context data platform for building observable, self-learning AI agents
Language modeling in a sentence representation space
Proofs, cases, concept supplements, and reference explanations
Di♪♪Rhythm: Blazingly Fast & Simple End-to-End Song Generation
Physical Symbolic Optimization
Synchronized Translation for Videos
Implementation of Video Diffusion Models