Alternatives to Merrymake
Compare Merrymake alternatives for your business or organization using the curated list below. SourceForge ranks the best alternatives to Merrymake in 2026. Compare features, ratings, user reviews, pricing, and more from Merrymake competitors and alternatives in order to make an informed decision for your business.
-
1
optivalue.ai
optivalue.ai
Sovereign AI that turns every answer into lasting expertise. Optivalue.ai helps large organizations build resilience in an increasingly complex environment: regulations piling up, audits tightening, senior expertise walking out the door. Our approach: • 89 Domain-Specific Language Models (DSLMs) — specialized by industry & function, not a generic LLM • Confidence score 0-100 on every answer • Ability to say "I don't know" rather than hallucinate • Strict sourcing: document + page + timestamp Our three pillars of trust: → Precision & epistemic lucidity (right answer, or no answer) → Traceability (every claim is defensible) → Confidentiality (your data never leaves — on-premise or sovereign cloud) At every interaction, your knowledge base grows. Expertise compounds. Your organization becomes more resilient. Trusted by: L'Oréal, Stellantis, Thales Alenia Space, Exaion (EDF Group), Equans, Mango Award: European Digital Sovereignty Prize 2026 (AI Category) -
2
Claude Opus 5
Anthropic
Claude Opus 5 is Anthropic’s advanced everyday AI model built for coding, knowledge work, problem-solving, visual outputs, and production AI workflows. The model delivers stronger performance than Opus 4.8 at the same base price and is positioned as a cost-effective alternative close to Claude Fable 5 frontier intelligence. Claude Opus 5 supports configurable effort settings so users can optimize for intelligence, speed, or token efficiency. It performs especially well on software engineering, automation, computer use, scientific research, and knowledge work evaluations. The model is available on Claude Max, Claude Pro, Claude API, Claude Code, and other Claude platforms, with Fast mode available at a higher price. Built for developers, researchers, enterprises, and everyday Claude users, Claude Opus 5 helps teams complete complex tasks with stronger verification, careful iteration, and practical cost efficiency.Starting Price: $5 per 1M tokens (input) -
3
Claude Mythos 5
Anthropic
Claude Mythos 5 is Anthropic’s most advanced restricted-access AI model, designed for trusted cyberdefenders, infrastructure providers, and select research organizations. It uses the same underlying model as Claude Fable 5 but provides lifted safeguards in approved areas for specialized high-trust use cases. The model delivers exceptional capabilities in cybersecurity, software engineering, scientific research, long-context reasoning, vision, and autonomous task execution. Anthropic initially deployed Claude Mythos 5 through Project Glasswing in collaboration with the U.S. government to help protect critical software and infrastructure. The model also shows strong potential in life sciences, including protein design, molecular biology hypothesis generation, and genomics research. Claude Mythos 5 is built for organizations that need frontier AI capabilities under controlled, trusted-access conditions.Starting Price: $10 per 1 million (input) -
4
Kimi K3
Moonshot AI
Kimi K3 is Moonshot AI’s most capable model, built for frontier intelligence scenarios such as software engineering, knowledge work, deep reasoning, and multimodal understanding. The model has 2.8 trillion parameters and uses Kimi Delta Attention, a hybrid linear attention mechanism, along with Attention Residuals for long-context performance. Kimi K3 supports a 1 million token context window, making it useful for analyzing large codebases, long documents, complex knowledge bases, and multi-step workflows. It includes native visual understanding for images and videos, with support for structured message formats, base64 image input, uploaded video files, and multimodal reasoning. Developers can use Kimi K3 through an OpenAI-compatible API with support for streaming, structured JSON output, partial mode, custom tools, dynamic tool loading, and automatic context caching.Starting Price: $3 per 1M tokens (input) -
5
Gemini 3.5 Flash
Google
Gemini 3.5 Flash is Google’s latest frontier AI model designed to combine advanced intelligence, high-speed performance, and agentic workflow execution for developers, enterprises, and everyday users. Built as part of the Gemini 3.5 family, the model excels at coding, long-horizon reasoning, multimodal understanding, and complex multi-step automation tasks while delivering significantly faster output speeds than many competing frontier models. Gemini 3.5 Flash powers AI agents capable of planning, executing, and managing workflows such as application development, codebase maintenance, data analysis, and financial document preparation through the Antigravity harness. The model also supports rich multimodal experiences by generating interactive graphics, dynamic web interfaces, animations, and advanced visual content. Gemini 3.5 Flash is integrated across Google products including the Gemini app, Google Search AI Mode, Google Antigravity, Google AI Studio, Android Studio, and more.Starting Price: $1.50 per 1M tokens (input) -
6
Sakana Fugu Ultra
Sakana AI
Sakana Fugu Ultra is the higher-performance version of Sakana Fugu, built to coordinate a deeper pool of expert AI agents for demanding, high-stakes tasks. The model operates through a single OpenAI-compatible API while dynamically orchestrating multiple powerful models behind the scenes. It is designed to maximize answer quality for complex workflows such as coding, code review, paper reproduction, cybersecurity analysis, scientific reasoning, patent investigation, and autonomous research. Fugu Ultra uses learned orchestration techniques to assemble, route, and coordinate agents instead of relying on hand-designed workflows or a single frontier model. Users can access advanced multi-agent intelligence without manually managing separate models, prompts, or collaboration patterns. Sakana Fugu Ultra is built for teams that need stronger performance, deeper reasoning, and more reliable results on difficult multi-step problems.Starting Price: $20 per month -
7
Fugu Cyber
Sakana AI
Fugu Cyber is a specialized multi-agent orchestration model purpose-built for modern cyber defense. It behaves like a single model through one API endpoint, but dynamically coordinates specialized agents to solve complex, multi-step security tasks without depending on one model provider. It focuses on two core defense workflows, analyzing complex codebases to verify real-world vulnerabilities and translating raw cyber threat intelligence into working detection rules. On CyberGym, which evaluates vulnerability analysis and verification, Fugu Cyber achieved an 86.9% success rate; on CTI-REALM, which measures detection-rule generation from threat reports, it reached 72.1%, placing it alongside leading cyber-focused frontier models. Fugu Cyber is intended to work as the reasoning engine inside broader security systems rather than as a standalone solution.Starting Price: $6 per 1M tokens (input) -
8
FLUX.2
Black Forest Labs
FLUX.2 is built for real production workflows, delivering high-quality visuals while maintaining character, product, and style consistency across multiple reference images. It handles structured prompts, brand-safe layouts, complex text rendering, and detailed logos with precision. The model supports multi-reference inputs, editing at up to 4 megapixels, and generates both photorealistic scenes and highly stylized compositions. With a focus on reliability, FLUX.2 processes real-world creative tasks—such as infographics, product shots, and UI mockups—with exceptional stability. It represents Black Forest Labs’ open-core approach, pairing frontier-level capability with open-weight models that invite experimentation. Across its variants, FLUX.2 provides flexible options for studios, developers, and researchers who need scalable, customizable visual intelligence. -
9
SWE-1.7
Cognition
SWE-1.7 is Cognition’s frontier software engineering model designed to deliver high intelligence at a lower rollout cost. The model is optimized for long-horizon agentic coding tasks, including debugging, feature implementation, codebase exploration, migrations, terminal workflows, and multilingual software engineering. SWE-1.7 was trained from a Kimi K2.7 base using large-scale reinforcement learning improvements across infrastructure, data quality, training stability, self-compaction, and long-running task execution. It is built to explore codebases thoroughly, probe edge cases, identify hidden requirements, and produce more complete end-to-end solutions. The model is available in Devin across web, desktop, and CLI through Cerebras at very high serving speeds. SWE-1.7 is positioned for developers and engineering teams that need cost-efficient frontier-level coding intelligence for complex real-world software work.Starting Price: $20/month -
10
Zyphra Cloud
Zyphra
Zyphra Cloud is a full-stack platform for open superintelligence, bringing advanced innovations from Zyphra Research into production for developers, enterprises, and frontier AI hyperscalers. It is designed for advanced AI systems with a focus on long-horizon agents, combining agent infrastructure, inference, agent environments, and compute into one unified platform for building and deploying open, sovereign AI at scale. Zyphra Cloud includes MAIA, a general open superagent for teams: a unified multimodal system that coordinates knowledge, communication, and execution across tools and workflows. MAIA is multiplayer by design, providing shared context, persistent memory, and coordinated execution across users and tools, while supporting interaction through language, audio, and vision in a single unified reasoning loop. Zyphra Inference is the first available component of the platform and is purpose-built to serve long-horizon agentic workloads. -
11
Mistral Compute
Mistral
Mistral Compute is a purpose-built AI infrastructure platform that delivers a private, integrated stack, GPUs, orchestration, APIs, products, and services, in any form factor, from bare-metal servers to fully managed PaaS. Designed to democratize frontier AI beyond a handful of providers, it empowers sovereigns, enterprises, and research institutions to architect, own, and optimize their entire AI environment, training, and serving any workload on tens of thousands of NVIDIA-powered GPUs using reference architectures managed by experts in high-performance computing. With support for region- and domain-specific efforts, defense technology, pharmaceutical discovery, financial markets, and more, it offers four years of operational lessons, built-in sustainability through decarbonized energy, and full compliance with stringent European data-sovereignty regulations. -
12
Step 3.5 Flash
StepFun
Step 3.5 Flash is an advanced open source foundation language model engineered for frontier reasoning and agentic capabilities with exceptional efficiency, built on a sparse Mixture of Experts (MoE) architecture that selectively activates only about 11 billion of its ~196 billion parameters per token to deliver high-density intelligence and real-time responsiveness. Its 3-way Multi-Token Prediction (MTP-3) enables generation throughput in the hundreds of tokens per second for complex multi-step reasoning chains and task execution, and it supports efficient long contexts with a hybrid sliding window attention approach that reduces computational overhead across large datasets or codebases. It demonstrates robust performance on benchmarks for reasoning, coding, and agentic tasks, rivaling or exceeding many larger proprietary models, and includes a scalable reinforcement learning framework for consistent self-improvement.Starting Price: Free -
13
Kimi K2
Moonshot AI
Kimi K2 is a state-of-the-art open source large language model series built on a mixture-of-experts (MoE) architecture, featuring 1 trillion total parameters and 32 billion activated parameters for task-specific efficiency. Trained with the Muon optimizer on over 15.5 trillion tokens and stabilized by MuonClip’s attention-logit clamping, it delivers exceptional performance in frontier knowledge, reasoning, mathematics, coding, and general agentic workflows. Moonshot AI provides two variants, Kimi-K2-Base for research-level fine-tuning and Kimi-K2-Instruct pre-trained for immediate chat and tool-driven interactions, enabling both custom development and drop-in agentic capabilities. Benchmarks show it outperforms leading open source peers and rivals top proprietary models in coding tasks and complex task breakdowns, while its 128 K-token context length, tool-calling API compatibility, and support for industry-standard inference engines.Starting Price: Free -
14
Harmonic Aristotle
Harmonic
Aristotle is the first AI model built from the ground up as a Mathematical Superintelligence (MSI), designed to deliver provably correct solutions to complex quantitative problems without hallucinations. When prompted with natural‑language math questions, it formalizes them in Lean 4, solves them via formally verified proofs, and returns both the proof and a natural‑language explanation. Unlike conventional language models that rely on probabilistic outputs, Aristotle’s MSI architecture replaces guesswork with provable logic, transparently flagging any errors or inconsistencies. The AI is accessible through a web interface and a developer API, enabling researchers to integrate its rigorous reasoning into workflows across fields such as theoretical physics, engineering, and computer science. -
15
STACKIT
STACKIT
STACKIT is a European cloud computing platform designed to provide scalable, secure, and data-sovereign cloud infrastructure for businesses, public institutions, and regulated industries. It delivers a full range of cloud services that allow organizations to run applications, store and process data, and build digital systems using infrastructure and platform tools hosted in European data centers. These services include infrastructure-as-a-service components such as virtual machines, storage, and networking, as well as platform-level services like managed databases, container runtimes, and application frameworks. STACKIT is built with a strong focus on digital sovereignty, meaning that data storage, processing, and operational control remain within the European Union and under European law, helping organizations meet strict data protection requirements such as GDPR. -
16
Qwen3.6-Max-Preview
Alibaba
Qwen3.6-Max-Preview is a next-generation frontier language model designed to push the limits of intelligence, instruction following, and real-world agent capabilities within the Qwen ecosystem. Building on the Qwen3 series, this preview release introduces stronger world knowledge, sharper instruction alignment, and significant improvements in agentic coding performance, enabling the model to better handle complex, multi-step tasks and software engineering workflows. It is engineered for advanced reasoning and execution scenarios, where the model not only generates responses but also interacts with tools, processes long contexts, and supports structured problem-solving across domains such as coding, research, and enterprise workflows. The architecture continues the Qwen focus on large-scale, high-efficiency models capable of handling extensive context windows and delivering consistent performance across multilingual and knowledge-intensive tasks.Starting Price: Free -
17
SWE-1.5
Cognition
SWE-1.5 is the latest agent-model release by Cognition, purpose-built for software engineering and characterized by a “frontier-size” architecture comprising hundreds of billions of parameters and optimized end-to-end (model, inference engine, and agent harness) for both speed and intelligence. It achieves near-state-of-the-art coding performance and sets a new benchmark in latency, delivering inference speeds up to 950 tokens/second, roughly six times faster than its predecessor Haiku 4.5 and thirteen times faster than Sonnet 4.5. The model was trained using extensive reinforcement learning in realistic coding-agent environments with multi-turn workflows, unit tests, quality rubrics, and browser-based agentic execution; it also benefits from tightly integrated software tooling and high-throughput hardware (including thousands of GB200 NVL72 chips and a custom hypervisor infrastructure). -
18
Mistral Large 3
Mistral AI
Mistral Large 3 is a next-generation, open multimodal AI model built with a powerful sparse Mixture-of-Experts architecture featuring 41B active parameters out of 675B total. Designed from scratch on NVIDIA H200 GPUs, it delivers frontier-level reasoning, multilingual performance, and advanced image understanding while remaining fully open-weight under the Apache 2.0 license. The model achieves top-tier results on modern instruction benchmarks, positioning it among the strongest permissively licensed foundation models available today. With native support across vLLM, TensorRT-LLM, and major cloud providers, Mistral Large 3 offers exceptional accessibility and performance efficiency. Its design enables enterprise-grade customization, letting teams fine-tune or adapt the model for domain-specific workflows and proprietary applications. Mistral Large 3 represents a major advancement in open AI, offering frontier intelligence without sacrificing transparency or control.Starting Price: Free -
19
GPT-Rosalind
OpenAI
GPT-Rosalind is a purpose-built frontier reasoning model developed by OpenAI to accelerate scientific research across biology, drug discovery, and translational medicine. It is designed specifically for life sciences workflows, where researchers must navigate large volumes of literature, experimental data, and specialized databases to generate and validate new ideas. It combines deep domain understanding in areas such as chemistry, genomics, protein engineering, and disease biology with advanced tool-use capabilities, allowing it to interact with scientific databases, analyze experimental outputs, and support complex, multi-step reasoning tasks. It can assist with evidence synthesis, hypothesis generation, literature review, sequence interpretation, and experimental planning, helping scientists move faster from raw data to actionable insights. GPT-Rosalind transforms complex, time-intensive research processes into more efficient AI-assisted workflows. -
20
Qwen3.7-Max
Alibaba
Qwen3.7-Max is Qwen’s latest proprietary model designed for the agent era, built to be a versatile agent foundation that is equally capable of writing and debugging code, automating office workflows, and sustaining autonomous browser sessions over long horizons. It reaches frontier-level coding performance, with stronger results across software engineering, terminal tasks, GUI grounding, web browsing, and agentic tool use. Qwen3.7-Max is designed to reduce the gap between model intelligence and real agent execution by supporting planning, long-context reasoning, reliable function calling, and multi-step task completion across complex workflows. It also strengthens multimodal and document-oriented work through Qwen Studio, which supports chatbot interaction, image and video understanding, image generation, document processing, presentation generation, coding assistance, deep research, and web development.Starting Price: Free -
21
LongCat-2.0
LongCat
LongCat-2.0 is a 1.6 trillion total-parameter Mixture-of-Experts language model built on AI ASIC superpods, with about 48 billion parameters activated per token and strong performance across coding and agentic tasks. It is a substantial step up from previous LongCat models, combining large-scale sparse architecture with dedicated post-training for real-world software engineering, tool use, long-context reasoning, and multi-step agent workflows. LongCat-2.0 is trained and deployed entirely on AI ASIC superpods, with pretraining spanning more than 35 trillion tokens and millions of accelerator-hours, demonstrating frontier-scale training on alternative hardware platforms. To strengthen long-horizon tasks, the model introduces LongCat Sparse Attention and is trained on hundreds of billions of tokens of 1M-context data, giving it native support for ultra-long context tasks and reliable long-document understanding. -
22
GPT-5.3-Codex
OpenAI
GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, designed to handle complex professional work on a computer. It combines frontier-level coding performance with advanced reasoning and real-world task execution. The model is faster than previous Codex versions and can manage long-running tasks involving research, tools, and deployment. GPT-5.3-Codex supports real-time interaction, allowing users to steer progress without losing context. It excels at software engineering, web development, and terminal-based workflows. Beyond code generation, it assists with debugging, documentation, testing, and analysis. GPT-5.3-Codex acts as an interactive collaborator rather than a single-turn coding tool. -
23
GPT-5.6 Sol Ultrafast
OpenAI
GPT-5.6 Sol Ultrafast is a new OpenAI API service tier that runs GPT-5.6 Sol up to 14× faster than Standard processing, bringing frontier intelligence to products and workflows where every second matters. Powered by Cerebras, it can generate up to 750 output tokens per second, allowing advanced reasoning to operate at real-time speeds without requiring a smaller or more specialized model. It is designed for time-sensitive business workflows where faster responses can change what AI can realistically do. Applications include incident response, where models can analyze logs, code changes, traces, and engineer reports while an outage is unfolding; financial research and security, where changing market signals and suspicious transactions can be assessed quickly; and customer support and voice, where complex issues can be resolved without interrupting a live conversation. In commerce, it can answer product questions, check inventory, and personalize recommendations. -
24
Arcee AI
Arcee AI
Arcee AI is a US-based open intelligence lab focused on building high-performance, open-weight AI models for developers and enterprises. It develops frontier AI systems designed for reasoning, scalability, and real-world applications. The company is known for its Trinity model family, which delivers advanced capabilities while remaining transparent and accessible. Arcee AI emphasizes continuous improvement through techniques like online reinforcement learning, allowing models to evolve after deployment. Its approach prioritizes cost efficiency, enabling powerful AI performance without excessive infrastructure costs. The platform supports developers with tools, APIs, and open-source resources to build intelligent applications. Overall, Arcee AI aims to make cutting-edge AI more accessible, practical, and scalable for a wide range of use cases. -
25
Celeris-1
Celeris-1
Celeris-1 is a low-latency, general-purpose language model platform and a diffusion model designed to deliver frontier-level intelligence at dramatically higher speed. Instead of generating one token at a time like traditional autoregressive models, Celeris uses a diffusion-based inference architecture that enables parallel generation and response times measured in milliseconds. On its published MMLU-Pro benchmark, Celeris-1 reaches 75.9 accuracy with a 158 ms median response time and 1,664 output tokens per second, placing it within a few points of frontier models while running more than 10x faster. The model is exposed through an OpenAI-compatible API, so developers can point existing SDKs and clients at Celeris with minimal code changes. Streaming is enabled for interactive applications, with responses as low as 24 ms and no buffering or batch delay.Starting Price: $0.20 per 1M tokens -
26
GPT-4V (Vision)
OpenAI
GPT-4 with vision (GPT-4V) enables users to instruct GPT-4 to analyze image inputs provided by the user, and is the latest capability we are making broadly available. Incorporating additional modalities (such as image inputs) into large language models (LLMs) is viewed by some as a key frontier in artificial intelligence research and development. Multimodal LLMs offer the possibility of expanding the impact of language-only systems with novel interfaces and capabilities, enabling them to solve new tasks and provide novel experiences for their users. In this system card, we analyze the safety properties of GPT-4V. Our work on safety for GPT-4V builds on the work done for GPT-4 and here we dive deeper into the evaluations, preparation, and mitigation work done specifically for image inputs. -
27
MiniMax M2.5
MiniMax
MiniMax M2.5 is a frontier AI model engineered for real-world productivity across coding, agentic workflows, search, and office tasks. Extensively trained with reinforcement learning in hundreds of thousands of real-world environments, it achieves state-of-the-art performance in benchmarks such as SWE-Bench Verified and BrowseComp. The model demonstrates strong architectural thinking, decomposing complex problems before generating code across more than ten programming languages. M2.5 operates at high throughput speeds of up to 100 tokens per second, enabling faster completion of multi-step tasks. It is optimized for efficient reasoning, reducing token usage and execution time compared to previous versions. With dramatically lower pricing than competing frontier models, it delivers powerful performance at minimal cost. Integrated into MiniMax Agent, M2.5 supports professional-grade office workflows, financial modeling, and autonomous task execution.Starting Price: Free -
28
ESMC
Biohub
ESMC is the latest in the ESM family of protein language models, establishing a new frontier in representation learning for protein biology. Trained on billions of evolutionary sequences, it learns representations that reflect a mechanistic reduction of protein structure and function. The model is built on a transformer architecture, supports sequences as its core modality, and is trained on up to 6 billion proteins. ESMC is designed for protein science research, including structure prediction, function annotation, protein design, and understanding evolutionary relationships between proteins. It can generate novel proteins from partial sequence, structure, or functional constraints, helping researchers explore new possibilities in protein design and biological discovery. The Biohub Platform provides access to ESMC through the API and the ESM Python package, with quickstart resources for installing the package, creating an API key, connecting to the platform.Starting Price: Free -
29
Odyssey
Odyssey ML
Odyssey is a frontier interactive video model that enables instant, real-time generation of video you can interact with. Just type a prompt, and the system begins streaming minutes of video that respond to your input. It shifts video from a static playback format to a dynamic, action-aware stream: the model is causal and autoregressive, generating each frame based solely on prior frames and your actions rather than a fixed timeline, enabling continuous adaptation of camera angles, scenery, characters, and events. The platform begins streaming video almost instantly, producing new frames every ~50 milliseconds (about 20 fps), so you don’t wait minutes for a clip, you engage in an evolving experience. Under the hood, the model is trained via a novel multi-stage pipeline to transition from fixed-clip generation to open-ended interactive video, allowing you to type or speak commands and explore an AI-imagined world that reacts in real time. -
30
Claude Haiku 4.5
Anthropic
Anthropic has launched Claude Haiku 4.5, its latest small-language model designed to deliver near-frontier performance at significantly lower cost. The model provides similar coding and reasoning quality as the company’s mid-tier Sonnet 4, yet it runs at roughly one-third of the cost and more than twice the speed. In benchmarks cited by Anthropic, Haiku 4.5 meets or exceeds Sonnet 4’s performance in key tasks such as code generation and multi-step “computer use” workflows. It is optimized for real-time, low-latency scenarios such as chat assistants, customer service agents, and pair-programming support. Haiku 4.5 is made available via the Claude API under the identifier “claude-haiku-4-5” and supports large-scale deployments where cost, responsiveness, and near-frontier intelligence matter. Claude Haiku 4.5 is available now on Claude Code and our apps. Its efficiency means you can accomplish more within your usage limits while maintaining premium model performance.Starting Price: $1 per million input tokens -
31
Claude Opus 3
Anthropic
Opus, our most intelligent model, outperforms its peers on most of the common evaluation benchmarks for AI systems, including undergraduate level expert knowledge (MMLU), graduate level expert reasoning (GPQA), basic mathematics (GSM8K), and more. It exhibits near-human levels of comprehension and fluency on complex tasks, leading the frontier of general intelligence. All Claude 3 models show increased capabilities in analysis and forecasting, nuanced content creation, code generation, and conversing in non-English languages like Spanish, Japanese, and French.Starting Price: Free -
32
OpenAI Astra
OpenAI
OpenAI Astra is an upcoming frontier AI model concept designed for advanced reasoning, multimodal understanding, agentic workflows, and long-horizon work. The model would be positioned to help users move from simple prompts to complete, high-quality deliverables across coding, research, business analysis, creative production, and knowledge work. Astra would combine text, image, document, voice, and tool-based capabilities into a more unified AI experience. It would be built for complex tasks that require planning, execution, verification, and iteration across multiple steps. Developers and teams could use Astra to power AI agents, productivity tools, coding assistants, research systems, and enterprise applications. As a next-generation OpenAI model concept, Astra would represent a shift toward AI systems that can reason, act, and collaborate across real-world workflows. -
33
Ling 3.0 Tiny
Ant Group
Ling 3.0 Tiny is an open-weights reasoning model with 7.9B total parameters, 1.3B active parameters, and a 262K-token context window. Built with a mixture-of-experts architecture, it extends the open-weights Pareto frontier for intelligence versus active parameters and is small enough to run locally in many settings. The model scores 25 on the Artificial Analysis Intelligence Index, comparable to gpt-oss-120b (high, 24) while using 15x fewer total parameters and 4x fewer active parameters. This parameter efficiency comes with relatively high token usage, with 213M output tokens required to run the Intelligence Index. Ling 3.0 Tiny also shows substantial improvements in hallucination behavior over Ling-mini-2.0, improving its AA-Omniscience score by 59 points while maintaining similar accuracy. Rather than guessing when uncertain, it attempted only 37% of questions in the evaluation, resulting in a 30% hallucination rate compared with 96% for the previous generation. -
34
Prime Intellect
Prime Intellect
Prime Intellect is the open superintelligence stack: an integrated compute, training, inference, and sandbox platform for teams that want to train, deploy, and continuously improve their own models. The stack is built around owning intelligence instead of waiting on frontier models to improve, giving users one loop for reinforcement learning environments, hosted evaluations, large-scale training, inference, and compute. In Lab, teams can post-train self-improving agents by turning tasks into RL environments, creating, developing, evaluating, and pushing them with the Prime CLI. The Environment Hub gives access to and contributions across 2,500+ open-source RL environments, while hosted evaluations let teams benchmark model performance across open-source models with no infrastructure or setup. Hosted Training supports large-scale models optimized for agentic workflows, managed training workflows with full visibility and control, and hands-on support from the applied research team. -
35
Phi-4
Microsoft
Phi-4 is a 14B parameter state-of-the-art small language model (SLM) that excels at complex reasoning in areas such as math, in addition to conventional language processing. Phi-4 is the latest member of our Phi family of small language models and demonstrates what’s possible as we continue to probe the boundaries of SLMs. Phi-4 is currently available on Azure AI Foundry under a Microsoft Research License Agreement (MSRLA) and will be available on Hugging Face. Phi-4 outperforms comparable and larger models on math related reasoning due to advancements throughout the processes, including the use of high-quality synthetic datasets, curation of high-quality organic data, and post-training innovations. Phi-4 continues to push the frontier of size vs quality. -
36
Fugu-Ultra v1.1
Sakana AI
Fugu-Ultra v1.1 is Sakana AI’s upgraded multi-agent orchestration model for complex coding, agentic work, and advanced reasoning. Rather than relying on one model, it dynamically coordinates a diverse pool of frontier models, selecting and combining specialized agents for each task while presenting the system through a single model interface. The v1.1 orchestration upgrade incorporates newer frontier models and improves performance across every tracked benchmark, with gains of up to 7.9 points over v1.0 and particularly strong results on ProgramBench and Terminal Bench 2.1. Fugu can now be used directly inside Claude Code through Claude Code-compatible endpoints, bringing a coordinated team of models into familiar terminal workflows for writing, debugging, reviewing, and executing code. A one-command installer configures the integration on Ubuntu and macOS, while manual setup is available for Windows and other environments.Starting Price: $6 per 1M tokens (input) -
37
Gemini 2.5 Deep Think
Google
Gemini 2.5 Deep Think is an enhanced reasoning mode within the Gemini 2.5 family that uses extended, parallel thinking and novel reinforcement learning techniques to tackle complex, multi-step problems in areas like math, coding, science, and strategic planning by generating and evaluating multiple lines of thought before responding, producing more detailed, creative, and accurate answers with support for longer replies and built-in tool integration (e.g., code execution and web search). Its performance shows state-of-the-art results on rigorous benchmarks, including LiveCodeBench V6 and Humanity’s Last Exam, and it demonstrates notable gains over previous versions in challenging domains, with internal evaluations also indicating improved content safety and tone-objectivity, though with a higher tendency to decline benign requests; Google is conducting frontier safety evaluations and implementing mitigations to manage risks as the model’s capabilities advance. -
38
Amazon Nova 2 Pro
Amazon
Amazon Nova 2 Pro is Amazon’s most advanced reasoning model, designed to handle highly complex, multimodal tasks across text, images, video, and speech with exceptional accuracy. It excels in deep problem-solving scenarios such as agentic coding, multi-document analysis, long-range planning, and advanced math. With benchmark performance equal or superior to leading models like Claude Sonnet 4.5, GPT-5.1, and Gemini Pro, Nova 2 Pro delivers top-tier intelligence across a wide range of enterprise workloads. The model includes built-in web grounding and code execution, ensuring responses remain factual, current, and contextually accurate. Nova 2 Pro can also serve as a “teacher model,” enabling knowledge distillation into smaller, purpose-built variants for specific domains. It is engineered for organizations that require precision, reliability, and frontier-level reasoning in mission-critical AI applications. -
39
OpenGPT-X
OpenGPT-X
OpenGPT-X is a German initiative focused on developing large AI language models tailored to European needs, emphasizing versatility, trustworthiness, multilingual capabilities, and open-source accessibility. The project brings together a consortium of partners to cover the entire generative AI value chain, from scalable, GPU-based infrastructure and data for training large language models to model design and practical applications through prototypes and proofs of concept. OpenGPT-X aims to advance cutting-edge research with a strong focus on business applications, thereby accelerating the adoption of generative AI in the German economy. The project also emphasizes responsible AI development, ensuring that the models are trustworthy and align with European values and regulations. The project provides resources such as the LLM Workbook, and a three-part reference guide with resources and examples to help users understand the key features of large AI language models.Starting Price: Free -
40
Axiomatic AI
Axiomatic AI
Axiomatic AI is an advanced artificial intelligence platform designed to accelerate scientific research and engineering workflows by combining generative AI with mathematical verification and physics-based reasoning. It is built around a concept called Axiomatic Intelligence, which integrates frontier AI models with formal logic and domain-specific world models to ensure that outputs are not only generated but also mathematically and physically validated. Unlike conventional AI systems that produce plausible answers without guarantees of correctness, Axiomatic AI uses verification systems that test results against formal specifications and engineering constraints before returning them to the user. This approach allows the platform to support mission-critical tasks in areas such as photonics, electronics, thermal engineering, mechanics, and signal analysis. -
41
Gemini 3 Flash
Google
Gemini 3 Flash is Google’s latest AI model built to deliver frontier intelligence with exceptional speed and efficiency. It combines Pro-level reasoning with Flash-level latency, making advanced AI more accessible and affordable. The model excels in complex reasoning, multimodal understanding, and agentic workflows while using fewer tokens for everyday tasks. Gemini 3 Flash is designed to scale across consumer apps, developer tools, and enterprise platforms. It supports rapid coding, data analysis, video understanding, and interactive application development. By balancing performance, cost, and speed, Gemini 3 Flash redefines what fast AI can achieve. -
42
Liquid AI
Liquid AI
Our goal at Liquid is to build the most capable AI systems to solve problems at every scale, such that users can build, access, and control their AI solutions. This is to ensure that AI will be meaningfully, reliably, and efficiently integrated at all enterprises. Long term, Liquid will create and deploy frontier-AI-powered solutions that are available to everyone. We build white-box models within a white-box organization. -
43
rready
rready
rready builds sovereign enterprise software for work and innovation: from the European alternative to Jira & Confluence to AI-native innovation management - trusted by organizations including BMW, Tetra Pak, and Swisscom. rready work – sovara: A sovereign European alternative to Jira and Confluence, designed to give enterprises full control over systems, data, workflows, and the evolution of their technology landscape. Built for regulated environments, it combines predictable pricing, strong governance, and operational continuity with a future-ready architecture, without compromising stability. rready innovate: AI-native solutions for idea management, intrapreneurship, and continuous improvement. From initial concept through to measurable business impact, the platform enables structured, scalable, and secure innovation across the organization. Our mission is to give enterprises full ownership of how they manage work and drive innovation, now and in the future. -
44
Solar Pro 2
Upstage AI
Solar Pro 2 is Upstage’s latest frontier‑scale large language model, designed to power complex tasks and agent‑like workflows across domains such as finance, healthcare, and legal. Packaged in a compact 31 billion‑parameter architecture, it delivers top‑tier multilingual performance, especially in Korean, where it outperforms much larger models on benchmarks like Ko‑MMLU, Hae‑Rae, and Ko‑IFEval, while also excelling in English and Japanese. Beyond superior language understanding and generation, Solar Pro 2 offers next‑level intelligence through an advanced Reasoning Mode that significantly boosts multi‑step task accuracy on challenges ranging from general reasoning (MMLU, MMLU‑Pro, HumanEval) to complex mathematics (Math500, AIME) and software engineering (SWE‑Bench Agentless), achieving problem‑solving efficiency comparable to or exceeding that of models twice its size. Enhanced tool‑use capabilities enable the model to interact seamlessly with external APIs and data sources.Starting Price: $0.1 per 1M tokens -
45
Ornith-1.0
DeepReinforce
Ornith-1.0 is a self-improving family of models built specially for agentic coding tasks. It spans the full spectrum from compact 9B Dense models suitable for edge device deployment to 397B MoE frontier-scale models optimized for maximum performance, with variants including 9B Dense, 31B Dense, 35B MoE, and 397B MoE. Built on top of pretrained Gemma 4 and Qwen 3.5, Ornith-1.0 achieves state-of-the-art performance among open-source models of comparable size on coding benchmarks. Its key innovation is a self-improving training framework that learns to generate both solution rollouts and the task-specific scaffolds that guide those rollouts. Instead of relying on fixed, human-designed harnesses, Ornith-1.0 treats the scaffold as a learnable object that co-evolves with the policy, allowing the model to jointly optimize the orchestration and the final solution.Starting Price: Free -
46
Marey
Moonvalley
Marey is Moonvalley’s foundational AI video model engineered for world-class cinematography, offering filmmakers precision, consistency, and fidelity across every frame. It is the first commercially safe video model, trained exclusively on licensed, high-resolution footage to eliminate legal gray areas and safeguard intellectual property. Designed in collaboration with AI researchers and professional directors, Marey mirrors real production workflows to deliver production-grade output free of visual noise and ready for final delivery. Its creative control suite includes Camera Control, transforming 2D scenes into manipulable 3D environments for cinematic moves; Motion Transfer, applying timing and energy from reference clips to new subjects; Trajectory Control, drawing exact paths for object movement without prompts or rerolls; Keyframing, generating smooth transitions between reference images on a timeline; Reference, defining appearance and interaction of individual elements.Starting Price: $14.99 per month -
47
EVI 3
Hume AI
Hume AI's EVI 3 is a third-generation speech-language model that streams in user speech and forms natural, expressive speech and language responses. At conversational latency, it produces the same quality of speech as our text-to-speech model, Octave. Simultaneously, it responds with the same intelligence as the most advanced LLMs of similar latency. It also communicates with reasoning models and web search systems as it speaks, “thinking fast and slow” to match the intelligence of any frontier AI system. EVI 3 can instantly generate new voices and personalities instead of being limited to a handful of speakers. For instance, users can speak to any of the more than 100,000 custom voices already created on our text-to-speech platform, each with an inferred personality. No matter the voice, it responds with a wide range of emotions or styles, implicitly or on command.Starting Price: Free -
48
Lambda
Lambda.ai
Lambda provides high-performance supercomputing infrastructure built specifically for training and deploying advanced AI systems at massive scale. Its Superintelligence Cloud integrates high-density power, liquid cooling, and state-of-the-art NVIDIA GPUs to deliver peak performance for demanding AI workloads. Teams can spin up individual GPU instances, deploy production-ready clusters, or operate full superclusters designed for secure, single-tenant use. Lambda’s architecture emphasizes security and reliability with shared-nothing designs, hardware-level isolation, and SOC 2 Type II compliance. Developers gain access to the world’s most advanced GPUs, including NVIDIA GB300 NVL72, HGX B300, HGX B200, and H200 systems. Whether testing prototypes or training frontier-scale models, Lambda offers the compute foundation required for superintelligence-level performance. -
49
Reactor
Reactor
Reactor is building the missing layer for world models and invites users to experience real-time world models through an early preview. Its product direction centers on worlds generated in real time, where pixels, sounds, and actions can be produced on the fly, changing how people interact with software and, eventually, the physical world. The preview is the first step toward that reality, letting users experience AI-generated worlds running on global low-latency infrastructure. Reactor’s work is focused on the next frontier of AI, real-time world models that people, agents, and robots can drive frame by frame. Rather than treating generated video as something passive to watch, Reactor points toward interactive environments that can be inhabited, controlled, and shaped as they generate. Its research and product focus includes real-time interactivity, inference, controllable world models, and systems that make dynamic visual environments responsive enough for live experiences.Starting Price: Free -
50
Xinity
Xinity
Xinity is open-source, OpenAI-compatible LLM inference software that lets European enterprises run generative AI entirely on their own servers. The platform installs on existing hardware and exposes an OpenAI-compatible API, so existing applications migrate by changing one base URL. No cloud dependency, no data egress, no exposure to the US CLOUD Act. The core engine is open source under Apache 2.0 and supports open-weight models, including European sovereign models, with automatic model routing, audit trails on every inference request, role-based access control, and multi-node orchestration. Xinity is built in Vienna, Austria for regulated industries such as finance, healthcare, legal, public sector, and media, including fully air-gapped environments, and is designed for GDPR and EU AI Act requirements.