Large Multimodal Models for Video Understanding and Editing
Faster and easier training and deployments
Official implementation of DreamCraft3D
Open-weight, large-scale hybrid-attention reasoning model
Let your agent control your phone
Deep Understanding AI Agents
Open source multi-agent RAG over a knowledge graph
Multilingual Document Layout Parsing in a Single Vision-Language Model
Shared repository for open-sourced projects from the Google AI Lang
Build multimodal AI applications with cloud-native stack
Ready-to-run cloud templates for RAG
Open-source industrial-grade ASR models
Fast-stable-diffusion + DreamBooth
A fast TTS architecture with conditional flow matching
Chinese and English multimodal conversational language model
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Stable Diffusion built-in to Blender
A Codex Skill for generating high-density, editable PowerPoints
AI Slack bot for reading, summarizing, and chatting with content
A Personalized LLM-powered Agent Frameworks
Running large language models on a single GPU
LongBench v2 and LongBench (ACL 25'&24')
LISA: Reasoning Segmentation via Large Language Model
Skywork-R1V is an advanced multimodal AI model series
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)