Debug, evaluate, and monitor your LLMapps, RAG systems, and agentic AI
Evaluate and compare LLM outputs, catch regressions, improve prompts
Experimental prompt optimization toolkit built around notebooks
A.S.E (AICGSecEval) is a repository-level AI-generated code security
Requirement-driven evaluation harness for AI agents and LLM
Visual tool for building, testing, and deploying AI agent workflows
AI CLI agent that writes code by iterating until tests pass
AI Agent Evaluator & Red Team Platform
Simple, unified interface to multiple Generative AI providers
Operating Layer for LabOS (Stanford-Princeton AI Co-Scientists)
Run Coding Agents in Sandboxes
Outcome driven agent development framework that evolves
Open-source MCP server that gives your coding agent
Skills, MCP servers, Custom Agents, Agents.md for SDKs
Distributed LLM and StableDiffusion inference
.NET Client for Telegram Bot API
Bench is a tool for evaluating LLMs for production use cases
Python example app from the OpenAI API quickstart tutorial
Supercharge your Playwright tests with AI
Ray Aviary - evaluate multiple LLMs easily
Testing framework that began with API and performance testing
Large dataset of coding contests designed for AI and ML model training
Pretrained models for TensorFlow.js
Coursera Machine Learning By Prof. Andrew Ng
A conversational agent prototyping platform