Document (PDF, Word, PPTX ...) extraction and parse API
Hypernetworks that adapt LLMs for specific benchmark tasks
Qwen-Image is a powerful image generation foundation model
A modular graph-based Retrieval-Augmented Generation (RAG) system
LLM abstractions that aren't obstructions
Simple, Pythonic building blocks to evaluate LLM applications
Qwen3-omni is a natively end-to-end, omni-modal LLM
Unifying 3D Mesh Generation with Language Models
A high-quality PDF to Markdown tool based on large language model
GLM-4-Voice | End-to-End Chinese-English Conversational Model
AI-powered code assistant for Vim. OpenAI and ChatGPT plugin for Vim
Knowledge Graph Generation from Any Text
Using AI models to automatically provide commentary and edit videos
LLM inference server with continuous batching & SSD caching
Toolkit for conversational AI
Data Infrastructure providing an approach to multimodal AI workloads
Build multimodal language agents for fast prototype and production
lightweight package to simplify LLM API calls
Search all of YouTube from the command line
Multilingual sentence & image embeddings with BERT
Capable of understanding text, audio, vision, video
Scalable data pre processing and curation toolkit for LLMs
Enhances Tesseract OCR output using LLMs (local or API)
Code and models for ICML 2024 paper, NExT-GPT
A list of free LLM inference resources accessible via API