Repo for SeedVR2 & SeedVR
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Industrial-level controllable zero-shot text-to-speech system
Long-form streaming TTS system for multi-speaker dialogue generation
General-purpose image editing model that delivers high-fidelity
Inference script for Oasis 500M
A theoretical reconstruction of the Claude Mythos architecture
Hunyuan Translation Model Version 1.5
Video understanding codebase from FAIR for reproducing video models
Sharp Monocular Metric Depth in Less Than a Second
A Production-ready Reinforcement Learning AI Agent Library
Open-source multi-speaker long-form text-to-speech model
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
Large Multimodal Models for Video Understanding and Editing
Open-Source Financial Large Language Models
C++ implementation of ChatGLM-6B & ChatGLM2-6B & ChatGLM3 & GLM4(V)
Advancing Open-source World Models
FAIR Sequence Modeling Toolkit 2
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
An AI-powered security review GitHub Action using Claude
The official PyTorch implementation of Google's Gemma models
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
GLM-4 series: Open Multilingual Multimodal Chat LMs