GalileoCisco
|
||||||
Related Products
|
||||||
About
AgentBench is an evaluation framework specifically designed to assess the capabilities and performance of autonomous AI agents. It provides a standardized set of benchmarks that test various aspects of an agent's behavior, such as task-solving ability, decision-making, adaptability, and interaction with simulated environments. By evaluating agents on tasks across different domains, AgentBench helps developers identify strengths and weaknesses in the agents’ performance, such as their ability to plan, reason, and learn from feedback. The framework offers insights into how well an agent can handle complex, real-world-like scenarios, making it useful for both research and practical development. Overall, AgentBench supports the iterative improvement of autonomous agents, ensuring they meet reliability and efficiency standards before wider application.
|
About
Galileo is an AI observability and eval engineering platform that helps teams measure, protect, and improve AI systems in development and production. Now part of Cisco, the platform turns offline evaluations into production guardrails so teams can stop AI failures instead of only monitoring them. Galileo helps users build datasets from synthetic, development, and live production data, while capturing subject matter expert annotations as ground truth. The platform includes more than 20 out-of-the-box evals for RAG, agents, safety, security, and custom evaluation workflows. Its Luna models distill expensive LLM-as-judge evaluators into lower-cost, low-latency guardrails that can monitor production traffic. Built for enterprises and developers, Galileo helps teams debug failures, improve agent behavior, enforce AI policies, and ship AI applications with more confidence.
|
|||||
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
|||||
Audience
AI developers wanting a tool to manage and evaluate their LLMs
|
Audience
Galileo is best suited for AI engineering teams, enterprise AI teams, developers, data scientists, ML engineers, platform teams, RAG builders, agent builders, safety teams, security teams, and organizations that need AI observability, eval engineering, production guardrails, ground-truth datasets, LLM-as-judge optimization, hallucination detection, agent debugging, safety evals, security evals, and reliable AI deployment
|
|||||
Support
Phone Support
24/7 Live Support
Online
|
Support
Phone Support
24/7 Live Support
Online
|
|||||
API
Offers API
|
API
Offers API
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
No information available.
Free Version
Free Trial
|
Pricing
No information available.
Free Version
Free Trial
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Webinars
Live Online
In Person
|
Training
Documentation
Webinars
Live Online
In Person
|
|||||
Company InformationAgentBench
China
llmbench.ai/agent
|
Company InformationCisco
United States
www.galileo.ai/
|
|||||
Alternatives |
Alternatives |
|||||
|
|
||||||
|
|
||||||
Categories |
Categories |
|||||
Integrations
Amazon Bedrock
Amazon SageMaker
Azure OpenAI Service
Gemini
Gemini 1.5 Flash
Gemini 1.5 Pro
Gemini 2.0
Gemini 2.0 Flash
Gemini Enterprise
Gemini Enterprise Agent Platform
|
Integrations
Amazon Bedrock
Amazon SageMaker
Azure OpenAI Service
Gemini
Gemini 1.5 Flash
Gemini 1.5 Pro
Gemini 2.0
Gemini 2.0 Flash
Gemini Enterprise
Gemini Enterprise Agent Platform
|
|||||
|
|
|