Bidirectional token-classification model for identifiable info
Long-form streaming TTS system for multi-speaker dialogue generation
This repository contains the official implementation of FastVLM
Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI
Research code artifacts for Code World Model (CWM)
An Efficient Agentic Model for Computer Use
LLM-based Reinforcement Learning audio edit model
1B text generation model based on the HRM architecture
Robust Speech Recognition Across Languages, Dialects
Codex plugin that turns attached object images into code-only
VMZ: Model Zoo for Video Modeling
Official implementation of Watermark Anything with Localized Messages
Video understanding codebase from FAIR for reproducing video models
CLIP, Predict the most relevant text snippet given an image
Qwen3-omni is a natively end-to-end, omni-modal LLM
Achieving 3+ generation speedup on reasoning tasks
Ultra-Efficient LLMs on End Device
Foundation Models for Time Series
A Production-ready Reinforcement Learning AI Agent Library
A PyTorch library for implementing flow matching algorithms
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
code for Mesh R-CNN, ICCV 2019
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
Language modeling in a sentence representation space
An AI-powered security review GitHub Action using Claude