OCR expert VLM powered by Hunyuan's native multimodal architecture
Collection of reference environments, offline reinforcement learning
LLM training in simple, raw C/CUDA
PPTAgent: Generating and Evaluating Presentations
A simple, secure MCP-to-OpenAPI proxy server
Implementation of "MobileCLIP" CVPR 2024
Code release for Cut and Learn for Unsupervised Object Detection
VMZ: Model Zoo for Video Modeling
Official implementation of Watermark Anything with Localized Messages
Training Large Language Model to Reason in a Continuous Latent Space
High-resolution models for human tasks
Code for the paper "Evaluating Large Language Models Trained on Code"
Tool for exploring and debugging transformer model behaviors
CLIP, Predict the most relevant text snippet given an image
Ling is a MoE LLM provided and open-sourced by InclusionAI
A Unified Framework for Text-to-3D and Image-to-3D Generation
Multimodal Diffusion with Representation Alignment
Personalize Any Characters with a Scalable Diffusion Transformer
Talk to Your AI Agents from Anywhere
The NVIDIA AgentIQ toolkit is an open-source library
Extensible AGI Framework
Open Source Generative Process Automation
SWE-agent takes a GitHub issue and tries to automatically fix it
Finding the Scaling Law of Agents. A multi-agent framework
PraisonAI application combines AutoGen and CrewAI or similar framework