RGBD video generation model conditioned on camera input
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training
Use ChatGPT to summarize the arXiv papers
Community plugin marketplace for Claude Cowork and Claude Code
AI PPT Track Terminator, the strongest PPT Skill ever
An Efficient Agentic Model for Computer Use
Audio foundation model excelling in audio understanding
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Infinite Worlds with Versatile Interactions
Tiny vision language model
The official PyTorch implementation of Google's Gemma models
Generates original ARC-AGI-1-style tasks distribution-matched
Codex plugin that turns attached object images into code-only
An Open Real-time Video-Language Interaction System
A 0.1B Omni model trained from scratch
26m function call model that runs on incredibly small devices
Open Source Speech Language Model
Open-source industrial-grade ASR models
Fast-stable-diffusion + DreamBooth
A Pragmatic VLA Foundation Model
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Multimodal embedding and reranking models built on Qwen3-VL
Implementation of "MobileCLIP" CVPR 2024
VMZ: Model Zoo for Video Modeling