Long-form streaming TTS system for multi-speaker dialogue generation
GPT4V-level open-source multi-modal model based on Llama3-8B
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
Contexts Optical Compression
An experimental version of DeepSeek model
Provides convenient access to the Anthropic REST API from any Python 3
A Powerful Native Multimodal Model for Image Generation
Open-Source Financial Large Language Models
GLM-4-Voice | End-to-End Chinese-English Conversational Model
Qwen2.5-VL is the multimodal large language model series
scikit-learn compatible tabular foundation model
New family of code large language models (LLMs)
Video understanding codebase from FAIR for reproducing video models
Tongyi Deep Research, the Leading Open-source Deep Research Agent
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model
Pokee Deep Research Model Open Source Repo
An AI-powered security review GitHub Action using Claude
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Multi-modal large language model designed for audio understanding
Large Multimodal Models for Video Understanding and Editing
Miso TTS is an 8 billion, highly emotive text-to-speech model
Qwen2.5-Coder is the code version of Qwen2.5, the large language model
Open Multilingual Multimodal Chat LMs
Towards Real-World Vision-Language Understanding
The ChatGPT Retrieval Plugin lets you easily find personal documents