Video understanding codebase from FAIR for reproducing video models
ICLR2024 Spotlight: curation/training code, metadata, distribution
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Language modeling in a sentence representation space
Stable Diffusion WebUI Forge is a platform on top of Stable Diffusion
High-Resolution Image Synthesis with Latent Diffusion Models
A Conversational Speech Generation Model
Di♪♪Rhythm: Blazingly Fast & Simple End-to-End Song Generation
Qwen2.5-Coder is the code version of Qwen2.5, the large language model
Open Multilingual Multimodal Chat LMs
ChatGLM-6B: An Open Bilingual Dialogue Language Model
Pushing the Limits of Mathematical Reasoning in Open Language Models
Towards Real-World Vision-Language Understanding
The ChatGPT Retrieval Plugin lets you easily find personal documents
Code release for "Masked-attention Mask Transformer
PyTorch implementation of MAE
Generate embeddings from large-scale graph-structured data
Efficient Image Captioning code in Torch, runs on GPU
Model that fuses instruct, reasoning and agentic skills
High-efficiency reasoning and agentic intelligence model
JetBrains’ 4B parameter code model for completions
OpenAI’s compact 20B open model for fast, agentic, and local use
CTC-based forced aligner for audio-text in 158 languages
Vision-language-action model for robot control via images and text