GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
High-Fidelity and Controllable Generation of Textured 3D Assets
Qwen2.5-VL is the multimodal large language model series
All-in-one WebUI for AI generative image and video creation
Universal LLM Deployment Engine with ML Compilation
New family of code large language models (LLMs)
Controllable & emotion-expressive zero-shot TTS
Generate blog articles from video or audio
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Renderer for the harmony response format to be used with gpt-oss
A Python library for audio
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
OpenRecall is a fully open-source, privacy-first alternative
Generates original ARC-AGI-1-style tasks distribution-matched
Codex plugin that turns attached object images into code-only
Agent skill: make LLMs write docs in ASD-STE100
Audio Language Models are Few-Shot Learners
Open source healthcare AI
AI tool for automating desktop tasks via natural language input
Python observability platform for tracing apps, metrics, and logs
Open source RAG framework for building scalable modular AI apps
An on-premises, OCR-free unstructured data extraction
An open-source, modern-design AI training tracking and visualization
Large Language Model Principles and Practice Tutorial from Scratch
Definitions for AI/ML tasks like dataset creation