A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
SOTA Open Source TTS
The SOTA Open-Source Browser Agent
A Model Context Protocol (MCP) Gateway & Registry
Open-sourced unified customization model
Renderer for the harmony response format to be used with gpt-oss
Natural language workflows for AI agents
Bring the notion of Model-as-a-Service to life
Outcome driven agent development framework that evolves
Towards Human-Sounding Speech
Bidirectional token-classification model for identifiable info
ComfyUI wrapper nodes for WanVideo and related models
Achieving 3+ generation speedup on reasoning tasks
Real-World Centric Foundation GUI Agents
UI-TARS-desktop version that can operate on your local personal device
Spark-TTS Inference Code
Foundation model for image generation
A Pragmatic VLA Foundation Model
Talk to Your AI Agents from Anywhere
Python package built to ease deep learning on graph
LLM-based Reinforcement Learning audio edit model
Agent S: an open agentic framework that uses computers like a human
MOSS‑TTS Family open‑source speech and sound generation model
Driving with Graph Visual Question Answering
Build cross-modal and multimodal applications on the cloud