Repo of Qwen2-Audio chat & pretrained large audio language model
Audio Language Models are Few-Shot Learners
Audio foundation model excelling in audio understanding
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Infinite Worlds with Versatile Interactions
The official PyTorch implementation of Google's Gemma models
Multi-modal large language model designed for audio understanding
Large Multimodal Models for Video Understanding and Editing
Official implementation of DreamCraft3D
Open-weight, large-scale hybrid-attention reasoning model
Open-source industrial-grade ASR models
Fast-stable-diffusion + DreamBooth
Chinese and English multimodal conversational language model
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
ICLR2024 Spotlight: curation/training code, metadata, distribution
Language modeling in a sentence representation space
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training
A Conversational Speech Generation Model
High-Resolution Image Synthesis with Latent Diffusion Models
Memory-efficient and performant finetuning of Mistral's models
AI-powered tool to quickly remove watermarks from images flawlessly
Towards Real-World Vision-Language Understanding
Chat & pretrained large vision language model
Chat & pretrained large audio language model proposed by Alibaba Cloud
Pushing the Limits of Mathematical Reasoning in Open Language Models