GLM-4-Voice | End-to-End Chinese-English Conversational Model
CogView4, CogView3-Plus and CogView3(ECCV 2024)
scikit-learn compatible tabular foundation model
Model export recipes, Python primitives, and Swift runtime utilities
Open Source Speech Language Model
Open-source industrial-grade ASR models
Ling is a MoE LLM provided and open-sourced by InclusionAI
A Unified Framework for Text-to-3D and Image-to-3D Generation
Multimodal Diffusion with Representation Alignment
Foundation Models for Time Series
Hackable and optimized Transformers building blocks
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Z80-μLM is a 2-bit quantized language model
Miso TTS is an 8 billion, highly emotive text-to-speech model
Designed for text embedding and ranking tasks
OCR expert VLM powered by Hunyuan's native multimodal architecture
Audio Language Models are Few-Shot Learners
1B text generation model based on the HRM architecture
Robust Speech Recognition Across Languages, Dialects
Repo for SeedVR2 & SeedVR
An Open Real-time Video-Language Interaction System
Fast-stable-diffusion + DreamBooth
Repo of Qwen2-Audio chat & pretrained large audio language model
Tongyi Deep Research, the Leading Open-source Deep Research Agent
Open Frontier Intelligence