A PyTorch library for implementing flow matching algorithms
Hackable and optimized Transformers building blocks
Official implementation of DreamCraft3D
Diffusion Transformer with Fine-Grained Chinese Understanding
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model
LLM-based Reinforcement Learning audio edit model
Reproduction of Poetiq's record-breaking submission to the ARC-AGI-1
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
Language modeling in a sentence representation space
An AI-powered security review GitHub Action using Claude
Generating Immersive, Explorable, and Interactive 3D Worlds
Implementation of the Surya Foundation Model for Heliophysics
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
High-Fidelity and Controllable Generation of Textured 3D Assets
Multi-modal large language model designed for audio understanding
State-of-the-art (SoTA) text-to-video pre-trained model
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Tooling for the Common Objects In 3D dataset
CodeGeeX: An Open Multilingual Code Generation Model (KDD 2023)
CodeGeeX2: A More Powerful Multilingual Code Generation Model
Memory-efficient and performant finetuning of Mistral's models
ChatGLM-6B: An Open Bilingual Dialogue Language Model