Persistent context and multi-instance coordination
Multimodal embedding and reranking models built on Qwen3-VL
"Big Model" trains a visual multimodal VLM with 26M parameters
Language Model Reinforcement Learning Environments frameworks
Collection of reference environments, offline reinforcement learning
A simple, secure MCP-to-OpenAPI proxy server
The most powerful Android RPA agent framework
Implementation of "MobileCLIP" CVPR 2024
Code release for Cut and Learn for Unsupervised Object Detection
Training Large Language Model to Reason in a Continuous Latent Space
High-resolution models for human tasks
Video understanding codebase from FAIR for reproducing video models
Tool for exploring and debugging transformer model behaviors
Ling is a MoE LLM provided and open-sourced by InclusionAI
Lightning-Fast RL for LLM Reasoning and Agents. Made Simple & Flexible
Multimodal-Driven Architecture for Customized Video Generation
PraisonAI application combines AutoGen and CrewAI or similar framework
Low-code framework for building custom LLMs, neural networks
Conditional GAN for generating synthetic tabular data
Determined, deep learning training platform
AI agents running research on single-GPU nanochat training
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
High-Fidelity and Controllable Generation of Textured 3D Assets
Large Multimodal Models for Video Understanding and Editing
An Open Source text-to-speech system built by inverting Whisper