Alternatives to SWE-2

Compare SWE-2 alternatives for your business or organization using the curated list below. SourceForge ranks the best alternatives to SWE-2 in 2026. Compare features, ratings, user reviews, pricing, and more from SWE-2 competitors and alternatives in order to make an informed decision for your business.

  • 1
    Devin Desktop

    Devin Desktop

    Cognition

    Devin Desktop (formerly Windsurf) is an AI-powered development environment that combines a full-featured IDE with advanced coding agents in a unified workspace. Formerly known as Windsurf, the platform enables developers to manage local and cloud-based AI agents, delegate tasks, review code, and ship software without leaving their editor. Developers can use multiple coding agents simultaneously to research, write, test, debug, and improve code while maintaining full visibility into every change. Devin Desktop includes features such as agent orchestration, shared workspaces, intelligent code completion, contextual code search, and integrated review tools. The platform supports a wide range of models, extensions, language servers, and MCP integrations, allowing teams to work with their preferred tools and workflows. Devin Desktop helps engineering teams accelerate software development, improve productivity, and manage AI-assisted coding at scale.
  • 2
    Claude Fable 5.1
    Claude Fable 5.1 is Anthropic’s advanced AI model for coding, knowledge work, research, and long-running agentic tasks. It is designed to improve on Claude Fable 5 with stronger performance across software engineering, scientific research, multidisciplinary reasoning, computer use, business workflows, and complex problem solving. The model can handle extended multi-step work, verify its own results, diagnose difficult software issues, and operate effectively across tool-heavy workflows. Anthropic also reduced cache-read pricing for Fable 5.1, lowering typical usage costs compared with Fable 5 and creating larger savings for highly agentic workloads. Fable 5.1 includes updated safeguards intended to reduce false positives while allowing more legitimate cybersecurity tasks such as vulnerability discovery for defensive purposes. The model is available through Claude products, the Claude API, Amazon Web Services, Google Cloud, and Microsoft Azure.
    Starting Price: $10 per 1M tokens (input)
  • 3
    Claude Mythos 5.1
    Claude Mythos 5.1 is Anthropic’s newest Mythos-class model, designed for advanced cybersecurity, biology, scientific research, coding, and long-running knowledge work. It is the same underlying model as Claude Fable 5.1 but uses different safeguards: Fable 5.1 is generally available, while Mythos 5.1 is restricted to trusted access programs with safeguards specifically designed for cybersecurity and life sciences research. The model sets a new performance frontier for agentic coding and demonstrates the strongest cyber capabilities of any Anthropic model released to date. In scientific research, Mythos 5.1 can work with specialized tools and complex workflows across molecular design, computational biology, and other technical domains. In Anthropic’s experiments, it designed high-affinity protein binders across multiple targets and achieved its strongest measured hit rate to date. It also optimized seven open-source protein and genomics deep learning models.
  • 4
    GPT-5.6 Sol
    GPT-5.6 Sol is a next-generation OpenAI model designed for advanced reasoning, coding, agentic workflows, biology analysis, cybersecurity support, and complex knowledge work. It is part of the GPT-5.6 model family alongside Terra and Luna, with Sol positioned as the flagship model for the most demanding tasks. The model introduces a new max reasoning effort for deeper thinking and an ultra mode that uses subagents to accelerate complex work beyond a single-agent approach. GPT-5.6 Sol shows strong performance in command-line coding workflows, long-horizon security tasks, genomics analysis, vulnerability research, debugging, patch development, and defensive testing. OpenAI pairs the model’s stronger capabilities with layered safeguards, real-time misuse classifiers, account-level review, automated red-teaming, and enterprise controls for sensitive workflows. GPT-5.6 Sol helps developers, enterprises, researchers, and security teams complete sophisticated technical work.
    Starting Price: $4 per 1M tokens (input)
  • 5
    GPT-6 Astra
    GPT-6 Astra is OpenAI’s frontier AI model for computer use, software engineering, scientific research, cybersecurity, browsing, and complex professional work. It combines advanced reasoning with agentic capabilities that allow it to navigate software, use tools, conduct research, manipulate data, troubleshoot systems, and complete multistep workflows. Astra is also designed to produce polished documents, spreadsheets, presentations, websites, applications, and other business or technical artifacts while following existing templates and organizational standards. In Codex, the model introduces improved long-running context management that can preserve notes and retrieve information from earlier context windows during extended software engineering tasks. OpenAI positions Astra as its most aligned model to date, with improvements in respecting task boundaries, interpreting user intent, communicating limitations, and avoiding unauthorized actions.
    Starting Price: $10 per 1M tokens (input)
  • 6
    Grok 4.6

    Grok 4.6

    SpaceXAI

    Grok 4.6 is an xAI model designed for long-running agents, ambitious interactive projects, visual work, coding, research, and knowledge workflows. The model builds on Grok 4.5 with stronger support for multi-step tasks that require sustained reasoning across codebases, information analysis, application development, and work artifact creation. Grok 4.6 can help turn broad product ideas into working first versions by researching domains, structuring applications, implementing core interactions, and refining results through feedback. It is trained across agentic tasks such as knowledge work, general coding, kernel optimization, web development, computer-aided design, and other technical environments. The model is available in Cursor, Grok Build, the xAI API, and partners such as OpenRouter, Vercel, and Cloudflare. Built for developers, builders, and teams working on complex projects, Grok 4.6 helps accelerate coding, agentic workflows, visual applications, and technical execution.
    Starting Price: $2 per 1M tokens (input)
  • 7
    Claude Opus 5

    Claude Opus 5

    Anthropic

    Claude Opus 5 is Anthropic’s advanced everyday AI model built for coding, knowledge work, problem-solving, visual outputs, and production AI workflows. The model delivers stronger performance than Opus 4.8 at the same base price and is positioned as a cost-effective alternative close to Claude Fable 5 frontier intelligence. Claude Opus 5 supports configurable effort settings so users can optimize for intelligence, speed, or token efficiency. It performs especially well on software engineering, automation, computer use, scientific research, and knowledge work evaluations. The model is available on Claude Max, Claude Pro, Claude API, Claude Code, and other Claude platforms, with Fast mode available at a higher price. Built for developers, researchers, enterprises, and everyday Claude users, Claude Opus 5 helps teams complete complex tasks with stronger verification, careful iteration, and practical cost efficiency.
    Starting Price: $5 per 1M tokens (input)
  • 8
    Claude Fable 5
    Claude Fable 5 is an advanced AI model from Anthropic designed to assist with software engineering, research, knowledge work, vision tasks, and complex reasoning. Built on the Mythos-class architecture, it delivers significantly improved performance across coding, analysis, and long-context workflows. The model can handle extended autonomous tasks while maintaining focus and consistency over large amounts of information. Claude Fable 5 integrates advanced reasoning, multimodal understanding, and memory capabilities to support professional and enterprise use cases. Anthropic has implemented specialized safeguards that automatically route certain high-risk cybersecurity, biology, chemistry, and model distillation requests to a different model. Claude Fable 5 helps organizations and professionals accelerate complex work while maintaining strong safety and governance controls.
    Starting Price: $10 per 1 million (input)
  • 9
    Claude Mythos 5
    Claude Mythos 5 is Anthropic’s most advanced restricted-access AI model, designed for trusted cyberdefenders, infrastructure providers, and select research organizations. It uses the same underlying model as Claude Fable 5 but provides lifted safeguards in approved areas for specialized high-trust use cases. The model delivers exceptional capabilities in cybersecurity, software engineering, scientific research, long-context reasoning, vision, and autonomous task execution. Anthropic initially deployed Claude Mythos 5 through Project Glasswing in collaboration with the U.S. government to help protect critical software and infrastructure. The model also shows strong potential in life sciences, including protein design, molecular biology hypothesis generation, and genomics research. Claude Mythos 5 is built for organizations that need frontier AI capabilities under controlled, trusted-access conditions.
    Starting Price: $10 per 1 million (input)
  • 10
    Claude Sonnet 5
    Claude Sonnet 5 is Anthropic's latest AI model, designed to deliver stronger agentic capabilities for coding, reasoning, tool use, and knowledge work while maintaining the efficiency of the Sonnet family. The model can independently plan tasks, use external tools such as browsers and terminals, and complete complex workflows that previously required larger AI models. Sonnet 5 significantly improves upon Claude Sonnet 4.6 with better reasoning, coding performance, reduced hallucinations, stronger safety behavior, and more effective autonomous task execution. It is available across Claude plans and through the Claude API with OpenAI-style developer access for application integration. Anthropic also introduced lower introductory API pricing, making Sonnet 5 a cost-effective option for developers building AI-powered products. By combining advanced agentic capabilities with improved safety and competitive pricing, Claude Sonnet 5 helps developers build more capable AI applications.
    Starting Price: $2 per 1M tokens (input)
  • 11
    GPT-5.6 Luna
    GPT-5.6 Luna is the fast and affordable model in OpenAI’s GPT-5.6 series, built to bring strong capability to users and developers who need practical intelligence with lower overhead. In the new GPT-5.6 naming system, the number identifies the model generation, while Sol, Terra, and Luna identify durable capability tiers that can advance on their own cadence, giving people and developers clearer choices across intelligence, speed, and cost. Luna sits alongside Sol, the flagship model, and Terra, the balanced model for everyday work, as part of a family designed for broader access to next-generation AI. During the limited preview, GPT-5.6 models are initially available through the API and Codex to a select group of trusted partners and organizations, with plans for broader availability in ChatGPT, Codex, and the API. OpenAI developed GPT-5.6 Sol, Terra, and Luna with its most robust safeguards to date, with configurations matched to each model’s capabilities.
    Starting Price: $0.20 per 1M tokens (input)
  • 12
    GPT-5.6 Terra
    GPT-5.6 Terra is a balanced model in the GPT-5.6 series designed for everyday work, coding, agentic workflows, cybersecurity support, biology analysis, and enterprise automation. It sits between GPT-5.6 Sol, the flagship model, and GPT-5.6 Luna, the faster and lower-cost option. Terra is positioned to deliver competitive performance to GPT-5.5 while being significantly cheaper to run. The model supports improved reasoning, coding, tool coordination, long-horizon workflows, and legitimate defensive security work. It is part of a model family built with layered safeguards, including trained refusals, real-time misuse classifiers, account-level review, differentiated access, monitoring, and continued red-team testing. GPT-5.6 Terra helps developers, enterprises, and technical teams access strong AI capabilities with a more practical balance of intelligence, speed, and cost.
    Starting Price: $2 per 1M tokens (input)
  • 13
    GLM-5.3
    GLM-5.3 is Z.ai’s frontier coding model designed for complex software engineering, long-horizon agent tasks, and advanced post-training research. The model uses the same base model as GLM-5.2, with improvements coming from scaled post-training across more environments, more diverse tasks, and larger compute investment. GLM-5.3 delivers stronger coding performance, better task ownership, improved benchmark results, and greater efficiency across realistic development workflows. It is built to handle complex coding tasks, production-style engineering work, research environments, automation tasks, and agentic workflows that require multi-step execution. The model also shows emergent cyber capabilities in vulnerability discovery and exploitation-chain reasoning, with safety evaluation and hardening planned before open-weight release.
  • 14
    Kimi K3

    Kimi K3

    Moonshot AI

    Kimi K3 is Moonshot AI’s most capable model, built for frontier intelligence scenarios such as software engineering, knowledge work, deep reasoning, and multimodal understanding. The model has 2.8 trillion parameters and uses Kimi Delta Attention, a hybrid linear attention mechanism, along with Attention Residuals for long-context performance. Kimi K3 supports a 1 million token context window, making it useful for analyzing large codebases, long documents, complex knowledge bases, and multi-step workflows. It includes native visual understanding for images and videos, with support for structured message formats, base64 image input, uploaded video files, and multimodal reasoning. Developers can use Kimi K3 through an OpenAI-compatible API with support for streaming, structured JSON output, partial mode, custom tools, dynamic tool loading, and automatic context caching.
    Starting Price: $3 per 1M tokens (input)
  • 15
    Qwen3.8-Max
    Qwen3.8-Max is Qwen’s most capable model to date, built as a Max-class AI model for coding, work, research, long-horizon tasks, and multimodal agents. It scales to 2.4 trillion parameters with 95 billion active parameters and is available through QwenCloud. The model is designed to complete complex, open-ended tasks end to end with greater reliability and minimal human involvement. Qwen3.8-Max supports autonomous coding workflows, agentic development, research reproduction, visual reasoning, document understanding, video analysis, and real-world productivity tasks. It can integrate with popular agent frameworks and coding assistants, including Claude Code, Codex, Qoder CLI, Qwen Code, and OpenClaw. Built for developers, researchers, enterprises, and AI agent builders, Qwen3.8-Max helps teams automate sophisticated work across code, documents, tools, interfaces, and multimodal content.
    Starting Price: $2 per 1M (input)
  • 16
    Grok 4.5

    Grok 4.5

    SpaceXAI

    Grok 4.5 is SpaceXAI’s advanced AI model built for coding, agentic tasks, engineering work, and knowledge-intensive productivity. The model is trained on coding, science, engineering, and math data, with reinforcement learning focused on multi-step software engineering and technical workflows. It is designed to handle real-world development tasks such as debugging, Rust and C/C++ work, terminal tasks, long-running agentic rollouts, and end-to-end app creation from a single prompt. Grok 4.5 is also built for fast serving, token efficiency, and lower-cost execution, with pricing based on input and output token usage. Beyond coding, the model supports business productivity tasks in Grok Build, including Excel modeling, PowerPoint diagram creation, Word writing, and research-assisted office workflows. Available through Grok Build, Cursor, and the SpaceXAI API console, Grok 4.5 gives developers and teams a high-performance model for building software, automating work, and more.
    Starting Price: $2 per million input tokens
  • 17
    Composer 2.5
    Composer 2.5 is the latest AI coding model released by Cursor, offering major improvements in intelligence, collaboration, and long-task performance compared to Composer 2. The model is designed to follow complex instructions more accurately while providing a smoother and more natural user experience during coding sessions. Cursor enhanced Composer 2.5 through larger-scale training, more advanced reinforcement learning environments, and improved behavioral tuning focused on communication and effort calibration. The model uses targeted reinforcement learning with textual feedback to correct specific mistakes during training, helping it avoid issues like invalid tool calls or poor coding behavior. Composer 2.5 was also trained using significantly more synthetic coding tasks, enabling it to handle increasingly difficult programming challenges and real-world development scenarios.
    Starting Price: $0.50/M input
  • 18
    GLM-5.3-Flash
    GLM-5.3-Flash is Z.ai’s natively multimodal model in the GLM-5 series (previously previewed as Ox Alpha), designed to deliver strong coding, agentic, visual, and knowledge-work performance at relatively low inference cost. It uses 320 billion total parameters with 18 billion active parameters, along with a hybrid architecture that combines sparse and linear attention to reduce the cost of long-context processing. The model supports context lengths of up to one million tokens and was trained on a 30-trillion-token multimodal corpus. GLM-5.3-Flash can reason across text, images, documents, interfaces, dashboards, and other visual information while using that feedback to refine its own outputs. Z.ai reports substantial gains over GLM-5.2 on coding and agentic benchmarks, including DeepSWE and AutomationBench, while approaching higher-cost frontier models on several evaluations.
    Starting Price: $0.15 per 1M tokens (input)
  • 19
    Nemotron 3 Ultra
    Nemotron 3 Nano is a compact, open large language model in NVIDIA’s Nemotron 3 family, designed for efficient agentic reasoning, conversational AI, and coding tasks. It uses a hybrid Mixture-of-Experts Mamba-Transformer architecture that activates only a small subset of parameters per token, enabling low-latency inference while maintaining strong accuracy and reasoning performance. It has approximately 31.6 billion total parameters with around 3.2 billion active (3.6 billion including embeddings), allowing it to achieve higher accuracy than previous Nemotron 2 Nano while using less computation per forward pass. Nemotron 3 Nano supports long-context processing of up to one million tokens, enabling it to handle large documents, multi-step workflows, and extended reasoning chains in a single pass. It is designed for high-throughput, real-time execution, excelling in multi-turn conversations, tool calling, and agent-based workflows where tasks require planning, reasoning, and more.
  • 20
    MiniMax M3

    MiniMax M3

    MiniMax

    MiniMax M3 is an open-weight multimodal AI model designed for coding, agentic workflows, long-context reasoning, and complex automation tasks. The model combines frontier-level coding performance, native multimodal understanding, and a context window of up to 1 million tokens. MiniMax M3 uses MiniMax Sparse Attention to improve long-context efficiency while reducing compute requirements for large-scale inputs. It supports text, image, and video understanding, making it useful for workflows that combine code, documents, visual references, and tool-driven tasks. The model is built for repository-scale reasoning, software engineering, autonomous task execution, tool calling, and multi-step agent workflows. MiniMax M3 helps developers, AI teams, and enterprises build capable agents that can reason across large contexts and work with multimodal information.
    Starting Price: $0.30 per million input tokens
  • 21
    Kimi K2.7 Code

    Kimi K2.7 Code

    Moonshot AI

    Kimi K2.7 Code is an open-source, coding-focused agentic AI model developed by Moonshot AI for long-horizon software engineering tasks. It is designed to improve coding performance, agent workflows, and real-world development assistance compared with earlier Kimi K2 versions. The model supports a 256K context window, making it useful for working with large codebases, long technical documents, and complex multi-step programming tasks. Kimi K2.7 Code is available through Kimi Code and API access, with OpenAI- and Anthropic-compatible options for easier integration into developer workflows. It is also listed on Hugging Face and supports deployment through inference engines such as vLLM, SGLang, and KTransformers. With improved agentic capabilities, long-context support, and reduced thinking-token usage compared with K2.6, Kimi K2.7 Code gives developers a flexible open-source option for AI-assisted coding.
  • 22
    Muse Spark 1.2
    Muse Spark 1.2 is Meta’s coding-focused model update designed to power Muse Code and improve software engineering workflows. The model is built for code generation, complex debugging, codebase understanding, long-horizon development tasks, and end-to-end developer workflows. Muse Spark 1.2 was co-trained with Muse Code to improve performance inside the terminal coding agent environment. It supports planning, goal conditioning, context compaction, subagent coordination, and iterative coding workflows across large repositories. The model was trained with expanded coding compute, diverse development environments, self-improvement loops, and long-running engineering tasks. Built for AI developers and software teams, Muse Spark 1.2 helps agents plan, write, validate, debug, and optimize code with greater autonomy.
    Starting Price: $1.25 per 1M tokens (input)
  • 23
    Laguna S 2.1
    Laguna S 2.1 is an open weight agentic coding model designed to pursue longer-horizon work and make effective use of reasoning. It uses a 118-billion-parameter Mixture-of-Experts architecture with 8 billion active parameters per token and supports a context window of up to one million tokens in both thinking and no-thinking modes. Its compact active size makes it suitable for complex work on local machines while remaining competitive with models many times larger on terminal, software-engineering, codebase-question-answering, and tool-use benchmarks. Laguna S 2.1 is built to keep working through difficult tasks with greater persistence, verification, and willingness to backtrack instead of declaring success too early. In demonstrated runs, it built and validated a browser rendering engine from an empty folder, optimized an agent harness for faster execution and substantially lower memory allocation, and completed extended mathematical research using the tools in its environment.
  • 24
    Seed2.1 Pro

    Seed2.1 Pro

    ByteDance

    Seed2.1 Pro is a next-generation AI productivity model built to handle complex, real-world work across general agents, code engineering, and multimodal understanding. It reliably executes multi-step tasks for high-value office work and everyday consultation, including project planning, file processing, research, tool use, spreadsheet analysis, lesson-plan slide generation, and industry report creation across tools and environments. In software development workflows, Seed2.1 Pro strengthens end-to-end delivery by improving requirement understanding, architecture design, coding, debugging, implementation, and validation. Its agent capabilities are designed to make steady progress on difficult tasks and return practical, verifiable results rather than isolated responses. The model also advances knowledge, reasoning, visual understanding, spatial reasoning, and long-context processing, giving agents a stronger foundation for complex decision-making and execution.
  • 25
    DeepSeek-V4-Flash
    DeepSeek-V4-Flash is a high-efficiency Mixture-of-Experts (MoE) language model designed for fast, scalable reasoning and text generation. It features 284 billion total parameters with 13 billion activated parameters, delivering strong performance while optimizing computational cost. The model supports an extensive context window of up to one million tokens, enabling it to process large documents and complex workflows with ease. Its hybrid attention architecture enhances long-context efficiency by reducing memory and compute requirements. Trained on over 32 trillion tokens, DeepSeek-V4-Flash demonstrates solid capabilities across knowledge, reasoning, and coding tasks. It is designed for scenarios where speed and efficiency are critical, offering a balance between performance and resource usage. The model also supports multiple reasoning modes, allowing users to adjust between faster outputs and deeper analysis.
    Starting Price: $0.14 per 1M tokens (input)
  • 26
    DeepSeek -V4.1-Flash
    DeepSeek-V4.1-Flash is a fast, versatile AI model designed for demanding coding, agentic, creative, and spatial reasoning workloads. Building on DeepSeek-V4-Flash, it emphasizes high-speed generation while maintaining strong performance on complex tasks, reaching more than 400 tokens per second and peaking at around 427 tokens per second in reported tests. The model can tackle advanced programming challenges, generate interactive 3D environments, create voxel-based designs, and reason about spatially complex scenes and simulations. Demonstrations include Minecraft-style worlds, classical Chinese gardens, racing environments, dungeon navigation, exploded camera views, and other applications requiring both code generation and an understanding of spatial relationships. Its capabilities make it suitable for rapid prototyping, game development, 3D workflows, architecture, research, and other technical or creative projects where iteration speed matters.
  • 27
    Grok 4.7

    Grok 4.7

    SpaceXAI

    Grok 4.7 is an upcoming xAI model expected to continue the Grok 4.x family’s focus on coding, reasoning, agentic workflows, and knowledge work. While xAI has not yet published an official Grok 4.7 launch page, model card, API slug, pricing, or benchmark report, the model is positioned as a future step beyond currently documented Grok 4-era releases. Grok 4.7 will likely build on xAI’s recent work around software engineering, tool use, long-context reasoning, multimodal capabilities, and real-time AI assistance. Developers and AI teams should treat Grok 4.7 as an upcoming model rather than a generally available product until xAI releases official documentation. Once available, it may be relevant for coding agents, research workflows, automation, technical support, and enterprise AI applications. Built for developers and power users tracking xAI’s roadmap, Grok 4.7 represents a likely next-stage model for advanced reasoning and agentic productivity.
  • 28
    Nemotron 3 Super
    Nemotron-3 Super is part of NVIDIA’s Nemotron 3 family of open models designed to enable advanced agentic AI systems that can reason, plan, and execute multi-step workflows across complex environments. The model introduces a hybrid Mamba-Transformer Mixture-of-Experts architecture that combines the efficiency of state-space Mamba layers with the contextual understanding of transformer attention, allowing it to process long sequences and complex reasoning tasks with high accuracy and throughput. This architecture activates only a subset of model parameters for each token, improving computational efficiency while maintaining strong reasoning capabilities and enabling scalable inference for large workloads. Nemotron-3 Super contains roughly 120 billion parameters with around 12 billion active during inference, accelerating multi-step reasoning and collaborative agent interactions across large contexts.
  • 29
    Nemotron 3.5 Lightning
    NVIDIA Nemotron 3.5 Lightning is an open 30B-parameter mixture-of-experts model with 3B active parameters, designed for high-volume, low-latency execution in long-running and always-on AI agents. Built for the execution layer of agentic systems, it handles frequent tasks such as tool calls, output validation, routine commands, and subagent delegation while larger reasoning models focus on planning and orchestration. Its MoE architecture activates only a fraction of parameters for each token, combining the capacity of a larger model with lower compute requirements. The model is trained for popular agent harnesses and supports speculative decoding through multi-token prediction, DFlash, and DSpark to improve inference speed across different serving scenarios. It is available with BF16 and NVFP4 checkpoints and can run from local systems such as DGX Spark and GeForce RTX hardware to data center environments.
  • 30
    Muse Spark 1.3
    Muse Spark 1.3 is an AI model with improved performance across agentic and coding tasks, designed to be smarter and more practical for real-world work. It sustains longer-horizon tasks by collaborating with users and managing multiple workflows in a single, long thread. Given an open-ended objective, it uses tools to build context across messy or conflicting sources, correct gaps in its plan, track what it has learned, and produce a final deliverable. It asks clarifying questions when prompts are ambiguous, requests help when stuck, and confirms before taking consequential actions. The model follows complex, long-form instructions more reliably, preserving detailed requirements across multi-step tasks without dropping constraints or drifting from the requested workflow. Improved multitasking allows it to map incoming prompts to the correct task even when users interrupt or redirect previous requests.
    Starting Price: $1.25 per 1M tokens (input)
  • 31
    Gemini 3.8 Flash
    Gemini 3.8 Flash is Google’s most intelligent Flash workhorse model, delivering significant improvements over 3.7 Flash across software engineering, agentic tasks, and critical multi-step reasoning in specialized domains. Built for long-horizon coding and autonomous agents, it can solve complex engineering problems end to end and delivers the dependability required for critical enterprise autonomy across specialized knowledge domains. The model shows stronger performance in quantitative and professional fields that require advanced analysis and reporting, as well as multi-step reasoning across STEM, humanities, and professional subjects. Its gains stem from a core design choice: Gemini 3.8 Flash works harder on complex tasks, executing additional reasoning steps and calling tools iteratively to maximize performance. At higher effort levels, it may use more tokens to pursue stronger results, while developers can select lower effort levels.
  • 32
    Gemini 4

    Gemini 4

    Google

    Gemini 4 is Google’s next-generation Gemini model family currently in development after the release of Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Google has confirmed that pre-training for Gemini 4 has begun, positioning it as the company’s most ambitious model training effort yet. The model is expected to advance Google’s frontier AI work across reasoning, coding, multimodal understanding, agentic workflows, and enterprise AI use cases. Because Gemini 4 has not been publicly released yet, official pricing, model cards, benchmarks, API details, and availability have not been published. Gemini 4 follows Google’s broader Gemini strategy of building models for developers, enterprises, consumer apps, and AI-powered products across Google’s ecosystem. Built for the next stage of AI agents and intelligent applications, Gemini 4 is likely to become a major foundation for future Google AI products once it becomes available.
  • 33
    Gemma 4

    Gemma 4

    Google

    Gemma 4 is an AI model introduced by Google and built on the Gemini architecture to deliver improved performance and flexibility. The model is designed to run efficiently on a single GPU or TPU, making it more accessible to developers and researchers. Gemma 4 enhances capabilities in natural language understanding and text generation, supporting a wide range of AI-driven applications. Its architecture allows it to handle complex tasks while maintaining efficient resource usage. Developers can use the model to build applications that rely on advanced language processing and automation. The design emphasizes scalability so that it can support both smaller projects and larger AI systems. By combining efficiency with powerful language capabilities, Gemma 4 helps advance the development of modern AI solutions.
  • 34
    Qwen3.8-2.4T-A95B
    Qwen3.8-2.4T-A95B is the largest open model in the Qwen3.8 family, bringing Qwen-Max-class capabilities to an open release. Built on the architectural foundation of Qwen3.5, it delivers substantial improvements across coding, professional work, research, and long-horizon agentic tasks, with a focus on carrying complex, multi-step work through to completion more reliably. The causal language model uses a mixture-of-experts architecture with 2.4 trillion total parameters and 95 billion activated parameters, including 512 experts with 10 routed and one shared expert active at a time. It supports a native context length of 262,144 tokens that can be extended to approximately 1.01 million tokens. Agent execution is strengthened through better autonomous planning and improved handling of environment feedback, while broader compatibility with popular agent harnesses and development tools simplifies integration into existing stacks.
  • 35
    Qwen3.8-27B
    Qwen3.8-27B is a compact open-weights model in Alibaba’s Qwen3.8 family, aimed at developers and researchers who want strong local AI performance without using the full Max-scale model. Reports from Alibaba’s Qwen3.8 launch state that Qwen3.8-27B was planned for open-weight release alongside Qwen3.8-Max, expanding access for builders working on AI applications. The model is positioned for coding, research, professional workflows, and local deployment scenarios where a 27B model can be more practical than frontier-scale systems. Qwen3.8’s broader launch emphasizes software development, document processing, data analysis, and professional “cowork” use cases. Qwen3.8-27B is especially relevant for teams that need a capable open model for experimentation, coding agents, assistant workflows, and self-hosted inference. Built for practical deployment, Qwen3.8-27B gives developers a smaller Qwen3.8 option for building AI tools, testing agents, and running advanced language model workflows.
  • 36
    Qwen3.8-Flash-Next
    Qwen3.8-Flash-Next is an open-weight multimodal Mixture-of-Experts model and an early preview of the architecture planned for Qwen4. It systematically upgrades attention, residual connections, embeddings, and optimization to improve capability, computational efficiency, model capacity, and training stability. Its hybrid architecture combines Gated DeltaNet, which efficiently compresses historical information, with Qwen Sparse Attention, which selects important context at the micro-block level to reduce attention and indexing costs on long sequences. Gated Residual widens the residual stream into four branches and dynamically controls information flow across layers, while N-gram Embedding adds large-scale local-pattern memory with very little extra per-token computation and can be offloaded to host memory. The model uses a 125B-parameter main network plus 51B N-gram embedding parameters, while activating only 6B parameters per token.
    Starting Price: $2 per 1M (input)
  • 37
    SWE-1.7

    SWE-1.7

    Cognition

    SWE-1.7 is Cognition’s frontier software engineering model designed to deliver high intelligence at a lower rollout cost. The model is optimized for long-horizon agentic coding tasks, including debugging, feature implementation, codebase exploration, migrations, terminal workflows, and multilingual software engineering. SWE-1.7 was trained from a Kimi K2.7 base using large-scale reinforcement learning improvements across infrastructure, data quality, training stability, self-compaction, and long-running task execution. It is built to explore codebases thoroughly, probe edge cases, identify hidden requirements, and produce more complete end-to-end solutions. The model is available in Devin across web, desktop, and CLI through Cerebras at very high serving speeds. SWE-1.7 is positioned for developers and engineering teams that need cost-efficient frontier-level coding intelligence for complex real-world software work.
    Starting Price: $20/month
  • 38
    SWE-1.6

    SWE-1.6

    Cognition

    SWE-1.6 is an engineering–focused AI model developed by Cognition and integrated into the Devin (Windsurf) environment, designed to optimize both raw intelligence and what the company calls “model UX,” or the overall feel and efficiency of interacting with an AI agent. It represents a new iteration in the SWE model family, improving performance on benchmarks such as SWE-Bench Pro by over 10% compared to SWE-1.5 while maintaining similar underlying capabilities. It was trained from scratch to jointly improve reasoning quality and user experience, addressing issues observed in earlier versions such as overthinking simple problems, taking too many steps, looping in repetitive reasoning, and relying excessively on terminal commands instead of specialized tools. SWE-1.6 introduces behavioral improvements such as more frequent parallel tool usage, faster context retrieval, and reduced need for user input, resulting in smoother and more efficient workflows.
  • 39
    Devin

    Devin

    Cognition AI

    Devin is an AI-driven software development assistant designed to collaborate with engineering teams to automate and accelerate coding tasks. It helps with tasks like setting up repositories, writing code, debugging, and performing migrations, all while working autonomously or alongside human developers. Devin is capable of learning from examples, making it more efficient over time. Its use has led to significant time and cost savings in large-scale projects, as seen in its deployment at Nubank, where it delivered 8-12x faster migrations and reduced costs by over 20x. Devin is particularly useful in refactoring and automating repetitive engineering tasks.
    Starting Price: $20/month
  • 40
    Lumen Outpost
    Lumen Outpost is Cosine’s targeted post-trained coding model, benchmarked against Kimi K2.6, its base model, GPT-5.5, GPT-5.4, and Gemini 3.1 Pro on highly complex, long-horizon coding tasks across 13 programming languages. The model is specialized not only for raw coding accuracy, but also for behavioral signals that matter in professional engineering workflows, including agent initiative, planning, scope discipline, action alignment, concise updates, and useful communication. Cosine’s benchmark report shows that highly targeted post-training transformed the base model’s capabilities, with Lumen Outpost outperforming Kimi K2.6 across Niche-Bench, Slop-Bench, Vibe-Bench, and cost per successful task. On Niche-Bench, an internal evaluation for niche, legacy, and environment-constrained programming languages, Lumen Outpost achieved a 53.9% score and led or tied in 9 of 13 assessed languages, with notable gains in Fortran, ABAP, Java, and Rust.
    Starting Price: $20 per month
  • 41
    GPT-5.3-Codex
    GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, designed to handle complex professional work on a computer. It combines frontier-level coding performance with advanced reasoning and real-world task execution. The model is faster than previous Codex versions and can manage long-running tasks involving research, tools, and deployment. GPT-5.3-Codex supports real-time interaction, allowing users to steer progress without losing context. It excels at software engineering, web development, and terminal-based workflows. Beyond code generation, it assists with debugging, documentation, testing, and analysis. GPT-5.3-Codex acts as an interactive collaborator rather than a single-turn coding tool.
  • 42
    GLM-5

    GLM-5

    Z.ai

    GLM-5 is Z.ai’s latest large language model built for complex systems engineering and long-horizon agentic tasks. It scales significantly beyond GLM-4.5, increasing total parameters and training data while integrating DeepSeek Sparse Attention to reduce deployment costs without sacrificing long-context capacity. The model combines enhanced pre-training with a new asynchronous reinforcement learning infrastructure called slime, improving training efficiency and post-training refinement. GLM-5 achieves best-in-class performance among open-source models across reasoning, coding, and agent benchmarks, narrowing the gap with leading frontier models. It ranks highly on evaluations such as Vending Bench 2, demonstrating strong long-term planning and operational capabilities. The model is open-sourced under the MIT License.
    Starting Price: Free
  • 43
    SWE-1.5

    SWE-1.5

    Cognition

    SWE-1.5 is the latest agent-model release by Cognition, purpose-built for software engineering and characterized by a “frontier-size” architecture comprising hundreds of billions of parameters and optimized end-to-end (model, inference engine, and agent harness) for both speed and intelligence. It achieves near-state-of-the-art coding performance and sets a new benchmark in latency, delivering inference speeds up to 950 tokens/second, roughly six times faster than its predecessor Haiku 4.5 and thirteen times faster than Sonnet 4.5. The model was trained using extensive reinforcement learning in realistic coding-agent environments with multi-turn workflows, unit tests, quality rubrics, and browser-based agentic execution; it also benefits from tightly integrated software tooling and high-throughput hardware (including thousands of GB200 NVL72 chips and a custom hypervisor infrastructure).
  • 44
    Kimi K2

    Kimi K2

    Moonshot AI

    Kimi K2 is a state-of-the-art open source large language model series built on a mixture-of-experts (MoE) architecture, featuring 1 trillion total parameters and 32 billion activated parameters for task-specific efficiency. Trained with the Muon optimizer on over 15.5 trillion tokens and stabilized by MuonClip’s attention-logit clamping, it delivers exceptional performance in frontier knowledge, reasoning, mathematics, coding, and general agentic workflows. Moonshot AI provides two variants, Kimi-K2-Base for research-level fine-tuning and Kimi-K2-Instruct pre-trained for immediate chat and tool-driven interactions, enabling both custom development and drop-in agentic capabilities. Benchmarks show it outperforms leading open source peers and rivals top proprietary models in coding tasks and complex task breakdowns, while its 128 K-token context length, tool-calling API compatibility, and support for industry-standard inference engines.
    Starting Price: Free
  • 45
    Laguna M.1

    Laguna M.1

    Poolside

    Laguna M.1 is Poolside’s most capable model for agentic coding, built and trained in-house for software development workflows. It is a 225B total-parameter Mixture of Experts model with 23B activated parameters, trained completely in-house on 30T tokens using 6,144 interconnected NVIDIA H200 GPUs. Poolside trained Laguna M.1 from scratch with its own data work, training codebase, and async on-policy reinforcement learning in its agent harness, all with agentic coding in mind. The model is designed to perform at its best inside Poolside’s coding agent, where it can reason through software tasks, interact with tools, edit code, run tests, and support longer autonomous development sessions. Laguna M.1 is built for developers and teams working on complex coding tasks that require stronger reasoning, architectural understanding, terminal use, and multi-step execution than lightweight models can provide.
    Starting Price: Free
  • 46
    Kimi K2.6

    Kimi K2.6

    Moonshot AI

    Kimi K2.6 is a next-generation agentic AI model developed by Moonshot AI, designed to push forward real-world execution, coding, and multi-step reasoning beyond earlier K2 and K2.5 versions. It builds on a Mixture-of-Experts architecture and the multimodal, agent-first foundation of the Kimi series, combining language understanding, coding, and tool use into a single system capable of planning and executing complex workflows. It introduces deeper reasoning capabilities and significantly improved agent planning, allowing it to break down tasks, coordinate tools, and handle multi-file or multi-step problems with greater accuracy and efficiency. It supports advanced tool calling with high reliability, enabling integration with external systems such as web search or APIs, and includes built-in validation mechanisms to ensure correct execution formats.
    Starting Price: Free
  • 47
    NVIDIA Llama Nemotron
    ​NVIDIA Llama Nemotron is a family of advanced language models optimized for reasoning and a diverse set of agentic AI tasks. These models excel in graduate-level scientific reasoning, advanced mathematics, coding, instruction following, and tool calls. Designed for deployment across various platforms, from data centers to PCs, they offer the flexibility to toggle reasoning capabilities on or off, reducing inference costs when deep reasoning isn't required. The Llama Nemotron family includes models tailored for different deployment needs. Built upon Llama models and enhanced by NVIDIA through post-training, these models demonstrate improved accuracy, up to 20% over base models, and optimized inference speeds, achieving up to five times the performance of other leading open reasoning models. This efficiency enables handling more complex reasoning tasks, enhances decision-making capabilities, and reduces operational costs for enterprises. ​
  • 48
    DeepSWE

    DeepSWE

    Agentica Project

    DeepSWE is a fully open source, state-of-the-art coding agent built on top of the Qwen3-32B foundation model and trained exclusively via reinforcement learning (RL), without supervised finetuning or distillation from proprietary models. It is developed using rLLM, Agentica’s open source RL framework for language agents. DeepSWE operates as an agent; it interacts with a simulated development environment (via the R2E-Gym environment) using a suite of tools (file editor, search, shell-execution, submit/finish), enabling it to navigate codebases, edit multiple files, compile/run tests, and iteratively produce patches or complete engineering tasks. DeepSWE exhibits emergent behaviors beyond simple code generation; when presented with bugs or feature requests, the agent reasons about edge cases, seeks existing tests in the repository, proposes patches, writes extra tests for regressions, and dynamically adjusts its “thinking” effort.
    Starting Price: Free
  • 49
    Inkling-Small

    Inkling-Small

    Thinking Machines Lab

    Inkling-Small is an efficient model that offers performance comparable to Inkling at a quarter of its size. It is a Mixture-of-Experts transformer with 276 billion total parameters and 12 billion active parameters, trained on NVIDIA GB300 NVL72 systems. It supports native reasoning across text, images, and audio, variable thinking effort, and context windows of up to one million tokens. Users adjust reasoning effort from minimal to extra high to balance performance and compute. Improved pre-training data, post-training with on-policy distillation from Inkling, and extended agentic coding reinforcement learning helped Inkling-Small surpass its larger counterpart on reasoning and coding benchmarks. It performs well in coding and tool-use harnesses, exceeds 80% on SWE-bench Verified, and combines strong reasoning with efficient output. Its encoder-free multimodal architecture processes audio as dMel spectrograms and images as 40-by-40-pixel patches alongside text tokens.
    Starting Price: $0.30 per million input tokens
  • 50
    Kimi K2.5

    Kimi K2.5

    Moonshot AI

    Kimi K2.5 is a next-generation multimodal AI model designed for advanced reasoning, coding, and visual understanding tasks. It features a native multimodal architecture that supports both text and visual inputs, enabling image and video comprehension alongside natural language processing. Kimi K2.5 delivers open-source state-of-the-art performance in agent workflows, software development, and general intelligence tasks. The model offers ultra-long context support with a 256K token window, making it suitable for large documents and complex conversations. It includes long-thinking capabilities that allow multi-step reasoning and tool invocation for solving challenging problems. Kimi K2.5 is fully compatible with the OpenAI API format, allowing developers to switch seamlessly with minimal changes. With strong performance, flexibility, and developer-focused tooling, Kimi K2.5 is built for production-grade AI applications.
    Starting Price: Free