Unified open dataset enabling cross-embodiment learning for robotics
Bring the notion of Model-as-a-Service to life
An orchestration platform for the development, production
A proxy server for multiple ollama instances with Key security
Real-Time Open-Ended Video Editing with Autoregressive Diffusion
Prompt Declaration Language is a declarative prompt programming lang
Pure Python FFmpeg-based live video / audio streaming to YouTube
Towards Human-Sounding Speech
Outcome driven agent development framework that evolves
Bidirectional token-classification model for identifiable info
ComfyUI wrapper nodes for WanVideo and related models
Achieving 3+ generation speedup on reasoning tasks
Real-World Centric Foundation GUI Agents
UI-TARS-desktop version that can operate on your local personal device
Harmonized and Coherent Human Image Animation
Spark-TTS Inference Code
Foundation model for image generation
A Pragmatic VLA Foundation Model
Talk to Your AI Agents from Anywhere
Python package built to ease deep learning on graph
LLM-based Reinforcement Learning audio edit model
Agent S: an open agentic framework that uses computers like a human
MOSS‑TTS Family open‑source speech and sound generation model
Driving with Graph Visual Question Answering
Build cross-modal and multimodal applications on the cloud