A modular graph-based Retrieval-Augmented Generation (RAG) system
Qwen3-omni is a natively end-to-end, omni-modal LLM
Simple, Pythonic building blocks to evaluate LLM applications
Accurate × Fast × Comprehensive
A Unified Framework for Text-to-3D and Image-to-3D Generation
Reading book source
Python framework for adversarial attacks, and data augmentation
A high-quality PDF to Markdown tool based on large language model
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
ComfyUI wrapper nodes for HunyuanVideo
Fast multimodal LLM for real-time voice interaction and AI apps
GLM-4-Voice | End-to-End Chinese-English Conversational Model
Unifying 3D Mesh Generation with Language Models
A Web UI for easy subtitle using whisper model
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
This does for Documents what repo-browser does for repos
Framework for building realtime multimodal voice AI agents apps
AI PPT Track Terminator, the strongest PPT Skill ever
Using AI models to automatically provide commentary and edit videos
Knowledge Graph Generation from Any Text
Han Language Processing
Open source personal AI Assistant for Linux, Windows and Mac
LLM inference server with continuous batching & SSD caching
State-of-the-art (SoTA) text-to-video pre-trained model
StreamSpeech is a seamless model for offline speech recognition