Fugu-Ultra v1.1

Fugu-Ultra v1.1

Sakana AI
+
+

Related Products

  • Gemini Enterprise Agent Platform
    985 Ratings
    Visit Website
  • LM-Kit.NET
    29 Ratings
    Visit Website
  • Dialpad Support
    1,588 Ratings
    Visit Website
  • Atera
    2,094 Ratings
    Visit Website
  • Creatio
    570 Ratings
    Visit Website
  • Sendbird
    165 Ratings
    Visit Website
  • Pipefy
    592 Ratings
    Visit Website
  • Docket
    59 Ratings
    Visit Website
  • NetBrain
    274 Ratings
    Visit Website
  • Checksum.ai
    1 Rating
    Visit Website

About

AgentBench is an evaluation framework specifically designed to assess the capabilities and performance of autonomous AI agents. It provides a standardized set of benchmarks that test various aspects of an agent's behavior, such as task-solving ability, decision-making, adaptability, and interaction with simulated environments. By evaluating agents on tasks across different domains, AgentBench helps developers identify strengths and weaknesses in the agents’ performance, such as their ability to plan, reason, and learn from feedback. The framework offers insights into how well an agent can handle complex, real-world-like scenarios, making it useful for both research and practical development. Overall, AgentBench supports the iterative improvement of autonomous agents, ensuring they meet reliability and efficiency standards before wider application.

About

Fugu-Ultra v1.1 is Sakana AI’s upgraded multi-agent orchestration model for complex coding, agentic work, and advanced reasoning. Rather than relying on one model, it dynamically coordinates a diverse pool of frontier models, selecting and combining specialized agents for each task while presenting the system through a single model interface. The v1.1 orchestration upgrade incorporates newer frontier models and improves performance across every tracked benchmark, with gains of up to 7.9 points over v1.0 and particularly strong results on ProgramBench and Terminal Bench 2.1. Fugu can now be used directly inside Claude Code through Claude Code-compatible endpoints, bringing a coordinated team of models into familiar terminal workflows for writing, debugging, reviewing, and executing code. A one-command installer configures the integration on Ubuntu and macOS, while manual setup is available for Windows and other environments.

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Audience

AI developers wanting a tool to manage and evaluate their LLMs

Audience

Research engineering teams that need multiple frontier models coordinated inside familiar terminal-based coding workflows

Support

Phone Support
24/7 Live Support
Online

Support

Phone Support
24/7 Live Support
Online

API

Offers API

API

Offers API

Screenshots and Videos

Screenshots and Videos

Pricing

No information available.
Free Version
Free Trial

Pricing

$6 per 1M tokens (input)
$6 per 1M tokens (input), $36 per 1M tokens (output)
Free Version
Free Trial

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Reviews/Ratings

Overall 5.0 / 5
ease 5.0 / 5
features 5.0 / 5

Pros & Cons from Real Users

Pros

  • Fugu-Ultra v1.1 is really interesting from a developer’s point of view because it feels like a different approach from the usual “one giant model does everything” setup. Instead, it acts more like an intelligent orchestration layer that can route work across multiple frontier models and agent patterns. That makes a lot of sense for coding and agentic workflows. Real development tasks are messy, and different parts of the job need different strengths: planning, repo search, debugging, terminal work, reasoning, code generation, and cleanup. I also like that Sakana added a Claude Code-compatible interface. Being able to use Fugu directly inside developer workflows makes it feel much more practical than something you only test in a browser or benchmark page. The benchmark gains are impressive too. If the reported improvements on coding and terminal tasks hold up in real projects, Fugu-Ultra v1.1 could be a strong option for developers building serious coding agents.

Cons

  • The main downside is that orchestration adds complexity. When a system is coordinating multiple models behind the scenes, I want strong transparency, logging, debugging tools, and predictable behavior before I trust it deeply in production. I would also want to test latency and cost carefully. Multi-agent systems can be powerful, but they can also become expensive or slow if they overthink simple tasks.

Training

Documentation
Webinars
Live Online
In Person

Training

Documentation
Webinars
Live Online
In Person

Company Information

AgentBench
China
llmbench.ai/agent

Company Information

Sakana AI
Founded: 2023
Japan
sakana.ai/fugu-1-1-claude-code-interface/

Alternatives

GLM-4.7

GLM-4.7

Zhipu AI

Alternatives

Claude Fable 5

Claude Fable 5

Anthropic
Claude Mythos 5

Claude Mythos 5

Anthropic
Claude Opus 5

Claude Opus 5

Anthropic
GLM-4.6

GLM-4.6

Zhipu AI
Sakana Fugu

Sakana Fugu

Sakana AI

Categories

Categories

Integrations

Sakana Fugu
Sakana Fugu Ultra

Integrations

Sakana Fugu
Sakana Fugu Ultra
Claim AgentBench and update features and information
Claim AgentBench and update features and information
Claim Fugu-Ultra v1.1 and update features and information
Claim Fugu-Ultra v1.1 and update features and information