Uncommon Objects in 3D dataset
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Language modeling in a sentence representation space
An AI-powered security review GitHub Action using Claude
GPT4V-level open-source multi-modal model based on Llama3-8B
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
Diversity-driven optimization and large-model reasoning ability
Repo of Qwen2-Audio chat & pretrained large audio language model
Multi-modal large language model designed for audio understanding
Open-source framework for intelligent speech interaction
Large Multimodal Models for Video Understanding and Editing
OCR expert VLM powered by Hunyuan's native multimodal architecture
Miso TTS is an 8 billion, highly emotive text-to-speech model
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Open-weight, large-scale hybrid-attention reasoning model
Large-language-model & vision-language-model based on Linear Attention
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training
LLM-based Reinforcement Learning audio edit model
Capable of understanding text, audio, vision, video
Chinese and English multimodal conversational language model
Tooling for the Common Objects In 3D dataset
Stable Diffusion WebUI Forge is a platform on top of Stable Diffusion
Memory-efficient and performant finetuning of Mistral's models
High-Resolution Image Synthesis with Latent Diffusion Models