Static Analyzer for Solidity
The library to build & auto-optimize LLM applications
Taming Stable Diffusion for Lip Sync
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Generating Immersive, Explorable, and Interactive 3D Worlds
Expressive Portrait Image Animation for Live Streaming
Open-source evaluation toolkit of large multi-modality models (LMMs)
Official implementation of Watermark Anything with Localized Messages
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
PDF to Markdown with vision models
Python inference and LoRA trainer package for the LTX-2 audio–video
Gracefully face hCaptcha challenge with multimodal llms
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Gemma open-weight LLM library, from Google DeepMind
PaddlePaddle End-to-End Development Toolkit
Powerful framework for controlling Android and iOS devices
Motion-controllable Video Generation via Latent Trajectory Guidance
Claude code for everything except coding
An AI-powered data science team of agents
Modular quant framework
ICLR2024 Spotlight: curation/training code, metadata, distribution
[CVPR 2025 Best Paper Award] VGGT
Unifying 3D Mesh Generation with Language Models
Flexible Photo Recrafting While Preserving Your Identity