OCR expert VLM powered by Hunyuan's native multimodal architecture
Official inference repo for FLUX.2 models
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
A Modular Simulation Framework and Benchmark for Robot Learning
Sparsity-aware deep learning inference runtime for CPUs
Ultralytics YOLO
A unified library of SOTA model optimization techniques
Automate native Android apps with AI using accessibility APIs
Open Source Computer Vision Library
Contexts Optical Compression
Workshop-Level Automated Scientific Discovery via Agentic Tree Search
A Foundation Model for Generalist Gaming Agents
Designed for training LLM/VLM agents via RL
Implementation of the Surya Foundation Model for Heliophysics
Welcome to GR00T Whole-Body Control (WBC)
LLM plugin providing access to models running on an Ollama server
Image processing in Python
Smart video converter using YOLOv8 and FFmpeg
Tiny vision language model
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
Data integration platform for ELT pipelines from APIs, databases
LISA: Reasoning Segmentation via Large Language Model
SpikingJelly is an open-source deep learning framework
Harmonized and Coherent Human Image Animation
Dataset Management Framework, a Python library and a CLI tool to build