One-ink editorial print image skill
Multimodal Agents as Smartphone Users, an LLM-based multimodal agent
1B text generation model based on the HRM architecture
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Generate Any 3D Scene in Seconds
A fast library for AutoML and tuning
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
Generate high-definition story short videos with one click using AI
Semi-Structured Agentic Framework. Workflows build themselves
Motion-controllable Video Generation via Latent Trajectory Guidance
A tool to use the Ai2 Open Coding Agents Soft-Verified Agents
Multimodal embedding and reranking models built on Qwen3-VL
Superfast AI decision making and processing of multi-modal data
PyTorch version of Stable Baselines
Diffusion Transformer with Fine-Grained Chinese Understanding
Context data platform for building observable, self-learning AI agents
SOTA discrete acoustic codec models with 40/75 tokens per second
Language modeling in a sentence representation space
Proofs, cases, concept supplements, and reference explanations
InvokeAI is a leading creative engine for Stable Diffusion models
UI-TARS-desktop version that can operate on your local personal device
Deep Learning Visualization Toolkit
High-Fidelity and Controllable Generation of Textured 3D Assets
Large Multimodal Models for Video Understanding and Editing
Integrate cutting-edge LLM technology quickly and easily into your app