LLM-based Reinforcement Learning audio edit model
High-Resolution Image Synthesis with Latent Diffusion Models
The most accurate natural language detection library for Python
A Multi-Modal World Model for Reconstructing, Generating, Simulation
Framework for building AI-powered interactive digital humans and agent
Toolkit for conversational AI
Running a 28.9M parameter LLM on an $8 microcontroller
ComfyUI wrapper nodes for WanVideo and related models
Han Language Processing
Paste Markdown and AI responses into Word Excel instantly fast
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Parse files for optimal RAG
Foundation model for image generation
Marrying Grounding DINO with Segment Anything & Stable Diffusion
HY-Motion model for 3D character animation generation
Open source NLP guide with models, methods, and real use cases
OCR expert VLM powered by Hunyuan's native multimodal architecture
High-Resolution 3D Assets Generation with Large Scale Diffusion Models
A Model Context Protocol (MCP) server
TextWorld is a sandbox learning environment for the training
MOSS‑TTS Family open‑source speech and sound generation model
Fast multimodal LLM for real-time voice interaction and AI apps
StarVector is a foundation model for SVG generation
ImageBind One Embedding Space to Bind Them All
tiktoken is a fast BPE tokeniser for use with OpenAI's models