Real-World Centric Foundation GUI Agents
Democratizing Reinforcement Learning for LLMs
Generate blog articles from video or audio
Provider-agnostic, open-source evaluation infrastructure
Reproduction of Poetiq's record-breaking submission to the ARC-AGI-1
Management of Yandex Station and other smart home devices
SOTA discrete acoustic codec models with 40/75 tokens per second
Controllable and fast Text-to-Speech for over 7000 languages
Stable Diffusion WebUI Forge is a platform on top of Stable Diffusion
Volcano Engine Reinforcement Learning for LLMs
An alignment auditing agent capable of exploring alignment hypothesis
Expose your FastAPI endpoints as Model Context Protocol (MCP) tools
FAIR Sequence Modeling Toolkit 2
Tooling for the Common Objects In 3D dataset
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
PyTorch code and models for VJEPA2 self-supervised learning from video
Language modeling in a sentence representation space
An AI-powered security review GitHub Action using Claude
GLM-4-Voice | End-to-End Chinese-English Conversational Model
GPT4V-level open-source multi-modal model based on Llama3-8B
An open sourced end-to-end VLM-based GUI Agent
A series of math-specific large language models of our Qwen2 series
Implementation of the Surya Foundation Model for Heliophysics