Collaborative & Open-Source Quality Assurance for all AI models
Evaluation suite designed to assess the performance of LLMs
Arcade Tool Development Kit (TDK), Worker, Evals, and CLI
AI agent harness for AI coding agents
Build high-quality LLM apps
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
Tools like web browser, computer access and code runner for LLMs
Lightweight framework for evaluating large language model performance
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
The open source post-building layer for agents
Open source platform for managing, testing, and deploying AI apps
Autonomous harness engineering
A high-performance ML model serving framework, offers dynamic batching
TextWorld is a sandbox learning environment for the training
Learn how to develop, deploy and iterate on production-grade ML
A framework that facilitates all stages of LLM development
ComfyUI wrapper nodes for WanVideo and related models
Leaderboard Comparing LLM Performance at Producing Hallucinations
General proxy performance testing tool based on Clash using Telegram
Reinforcement Learning for Humanoid Robot with Zero-Shot Sim2Real
Trainable, memory-efficient, and GPU-friendly PyTorch reproduction
A Conversational Speech Generation Model
Code repo for "WebArena to build Autonomous Agents
The most simple, flexible, and comprehensive OpenAI Gym trading
Hypergraph Transformer for Skeleton-based Action Recognition