Large Audio Language Model built for natural interactions
Stable Diffusion web UI
A lightweight text-to-speech model with zero-shot voice cloning
Follow along with my AI Agents Masterclass videos
Open source AI Agents hosted on the oTTomator Live Agent Studio
Context engineering is the new vibe coding
Chinese XLNet pre-trained model
Document Image Parsing via Heterogeneous Anchor Prompting”
Framework for building neural networks
The best ChatGPT that $100 can buy
4M: Massively Multimodal Masked Modeling
Guiding Instruction-based Image Editing via Multimodal Large Language
This repository contains the official implementation of FastVLM
Refer and Ground Anything Anywhere at Any Granularity
Supercharge Your LLM with the Fastest KV Cache Layer
Set of tools to assess and improve LLM security
ICLR2024 Spotlight: curation/training code, metadata, distribution
PyTorch code and models for V-JEPA self-supervised learning from video
A PyTorch library for implementing flow matching algorithms
An implementation of a deep learning recommendation model (DLRM)
Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI
Research code artifacts for Code World Model (CWM)
Diffusion Transformer with Fine-Grained Chinese Understanding
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model