Supercharge Your LLM Application Evaluations
CodiumAI Cover-Agent: An AI-Powered Tool for Automated Test Generation
Test-Time Reinforcement Learning
AI tool that generates tests to improve code coverage quickly
Collaborative & Open-Source Quality Assurance for all AI models
The easiest way to use deep metric learning in your application
Simple, unified interface to multiple Generative AI providers
General proxy performance testing tool based on Clash using Telegram
A python library that makes AMR parsing, generation and visualization
A powerful tool for automated LLM fuzzing
PaddlePaddle End-to-End Development Toolkit
Free, open source crypto trading bot
Requirement-driven evaluation harness for AI agents and LLM
Visual tool for building, testing, and deploying AI agent workflows
SWE-agent takes a GitHub issue and tries to automatically fix it
Evaluate and monitor ML models from validation to production
AI Agent Evaluator & Red Team Platform
Convert TensorFlow, Keras, Tensorflow.js and Tflite models to ONNX
Test Suites for validating ML models & data
Official inference library for Mistral models
Python library for portfolio optimization built on top of scikit-learn
Python SDK for agent monitoring, LLM cost tracking, benchmarking, etc.
MTEB: Massive Text Embedding Benchmark
Tools like web browser, computer access and code runner for LLMs
A benchmark built to evaluate and improve agent capabilities