Diversity-driven optimization and large-model reasoning ability
Pokee Deep Research Model Open Source Repo
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training
An Open Real-time Video-Language Interaction System
Audio Language Models are Few-Shot Learners
A 0.1B Omni model trained from scratch
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Stable Virtual Camera: Generative View Synthesis with Diffusion Models
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Multi-modal large language model designed for audio understanding
State-of-the-art (SoTA) text-to-video pre-trained model
Large Multimodal Models for Video Understanding and Editing
The official PyTorch implementation of Google's Gemma models
Long-form streaming TTS system for multi-speaker dialogue generation
General-purpose image editing model that delivers high-fidelity
Use ChatGPT to summarize the arXiv papers
Community plugin marketplace for Claude Cowork and Claude Code
Codex plugin that turns attached object images into code-only
An Efficient Agentic Model for Computer Use
New family of code large language models (LLMs)
Multimodal embedding and reranking models built on Qwen3-VL
LLM-based Reinforcement Learning audio edit model
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Infinite Worlds with Versatile Interactions
Tiny vision language model