Large Audio Language Model built for natural interactions
StreamSpeech is a seamless model for offline speech recognition
Interface for OuteTTS models
Omnilingual ASR Open-Source Multilingual SpeechRecognition
Self-supervised visual learning using momentum contrast in PyTorch
One-click local MCP server installation in desktop apps
A neural network that transforms a design mock-up into static websites
Diffusion Transformer with Fine-Grained Chinese Understanding
A solution to build and deploy MCP agents and applications
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model
A research prototype of a human-centered web agent
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
20+ high-performance LLMs with recipes to pretrain, finetune at scale
SAPIEN Manipulation Skill Framework
The fastest way to bring multi-agent workflows to production
CRAB: Cross-environment Agent Benchmark for Multimodal Language Model
The Memory layer for AI Agents
Module for automatic summarization of text documents and HTML pages
Supercharge Your LLM Application Evaluations
A Telegram bot that integrates with OpenAI's official ChatGPT APIs
A lightweight 3D Morphable Face Model library in modern C++
Data manipulation and transformation for audio signal processing
Open source libraries and APIs to build custom preprocessing pipelines
MII makes low-latency and high-throughput inference possible
A batteries-included library for building AI-powered software