Fast stable diffusion on CPU and AI PC
AlphaFold 3 inference pipeline
Awesome multilingual OCR toolkits based on PaddlePaddle
From Images to High-Fidelity 3D Assets
Industrial-level controllable zero-shot text-to-speech system
Tongyi Deep Research, the Leading Open-source Deep Research Agent
Generating Immersive, Explorable, and Interactive 3D Worlds
Native and Compact Structured Latents for 3D Generation
An Open Real-time Video-Language Interaction System
Advanced language and coding AI model
Video Object and Interaction Deletion
Robust Speech Recognition Across Languages, Dialects
1B text generation model based on the HRM architecture
code for Mesh R-CNN, ICCV 2019
Controllable & emotion-expressive zero-shot TTS
Pokee Deep Research Model Open Source Repo
A Multi-Modal World Model for Reconstructing, Generating, Simulation
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training
AI cognitive-enhancement Skills based on Anthropic's J-space
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Qwen3-omni is a natively end-to-end, omni-modal LLM
Project Lyra: Open Generative 3D World Models
Inference script for Oasis 500M
General-purpose image editing model that delivers high-fidelity