Controllable & emotion-expressive zero-shot TTS
A Systematic Framework for Interactive World Modeling
Reproduction of Poetiq's record-breaking submission to the ARC-AGI-1
Pokee Deep Research Model Open Source Repo
Uncommon Objects in 3D dataset
Language modeling in a sentence representation space
An AI-powered security review GitHub Action using Claude
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
Diversity-driven optimization and large-model reasoning ability
Repo of Qwen2-Audio chat & pretrained large audio language model
Multi-modal large language model designed for audio understanding
Open-source framework for intelligent speech interaction
Large Multimodal Models for Video Understanding and Editing
OCR expert VLM powered by Hunyuan's native multimodal architecture
Miso TTS is an 8 billion, highly emotive text-to-speech model
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Open-weight, large-scale hybrid-attention reasoning model
Large-language-model & vision-language-model based on Linear Attention
LLM-based Reinforcement Learning audio edit model
Capable of understanding text, audio, vision, video
Chinese and English multimodal conversational language model
Tooling for the Common Objects In 3D dataset
Stable Diffusion WebUI Forge is a platform on top of Stable Diffusion
High-Resolution Image Synthesis with Latent Diffusion Models
CodeGeeX: An Open Multilingual Code Generation Model (KDD 2023)