Debug, evaluate, and monitor your LLMapps, RAG systems, and agentic AI
Evaluate and compare LLM outputs, catch regressions, improve prompts
Experimental prompt optimization toolkit built around notebooks
A.S.E (AICGSecEval) is a repository-level AI-generated code security
AI CLI agent that writes code by iterating until tests pass
Requirement-driven evaluation harness for AI agents and LLM
Visual tool for building, testing, and deploying AI agent workflows
AI Agent Evaluator & Red Team Platform
Simple, unified interface to multiple Generative AI providers
Operating Layer for LabOS (Stanford-Princeton AI Co-Scientists)
Run Coding Agents in Sandboxes
Outcome driven agent development framework that evolves
Open-source MCP server that gives your coding agent
Skills, MCP servers, Custom Agents, Agents.md for SDKs
Distributed LLM and StableDiffusion inference
.NET Client for Telegram Bot API
Bench is a tool for evaluating LLMs for production use cases
Python example app from the OpenAI API quickstart tutorial
Supercharge your Playwright tests with AI
Ray Aviary - evaluate multiple LLMs easily
Testing framework that began with API and performance testing
Pretrained models for TensorFlow.js
Coursera Machine Learning By Prof. Andrew Ng
A conversational agent prototyping platform
A Free and Open Source Java Framework for Multiobjective Optimization