READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Cognition's New SWE-2 Coding AI Claims 64% Cost Edge Over Rival, By Its Own Benchmark

Cognition's New SWE-2 Coding AI Claims 64% Cost Edge Over Rival, By Its Own Benchmark
Cognition says its new SWE-2 model, launched September 10 in the Devin coding agent, matches Anthropic's top model on Cognition's own benchmark while costing 64% less to run. The numbers come entirely from Cognition, on a test Cognition built, so treat the win as a company claim, not an independent verdict.

Cognition released SWE-2 on Thursday, September 10, 2026, built into its Devin Desktop and CLI coding agent. 7, released in July: that model would read deep into a codebase before touching a single file, making its first actual code edit at a median of 48 steps into a task. SWE-2, running at its "medium" effort setting, makes that first edit at step 18, according to Cognition. That's a 62.5% cut in the exploration phase before real work starts. Cognition attributes the change to a new training method it calls a "Pareto-informed penalty algorithm." In plain terms: during reinforcement learning, the model got docked points every time it burned inference spend it didn't need. Cognition calls the resulting behavior "focused exploration" — the model learning which files in a codebase actually matter for a given task and skipping the rest.

The Numbers

According to Cognition On Cognition's own FrontierCode 1.1 Main benchmark, the company reports SWE-2 medium scores higher than SWE-1.7 while using 58% fewer turns and running at 81% lower average cost per task. Cognition also pits SWE-2 against what it identifies as the current FrontierCode 1.1 Main leader, a model it calls Anthropic's "Fable 5.1," which scores 50.9% on that benchmark. SWE-2 scores 50.0% — essentially a tie — while Cognition says it runs 64% cheaper. For enterprise customers already paying frontier-model rates to run Devin at scale, that's the pitch: near-identical benchmark performance at roughly a third of the line-item cost.

What's Missing From the Claim

Every number in this story comes from Cognition, measured on a benchmark Cognition built and controls. FrontierCode 1.1 Main is not a third-party industry standard like SWE-bench or HumanEval that independent researchers can audit freely. It's Cognition's in-house yardstick, which means Cognition is grading its own homework and declaring a near-tie against a competitor using its own scoring system. Companies test their own products constantly, and internal benchmarks are standard practice across the AI industry. But a 64%-cheaper claim measured on a proprietary test, with no independent lab or third-party evaluator cited, is a company's marketing number until someone outside Cognition replicates it. Enterprise buyers deciding whether to switch workloads should ask for that independent verification before taking the cost savings at face value. Cognition does offer some texture beyond the raw scores. The company says SWE-2 writes stronger test coverage, catching regressions rather than just satisfying the literal task specification. In one internal example Cognition describes, when an MCP integration wasn't available, SWE-2 reconstructed the missing data from Slack channel history already sitting in its context window rather than stalling out. And when challenged on a conclusion, Cognition says the model re-runs code to gather fresh evidence rather than just repeating its earlier answer. Those are the kinds of behavioral claims that matter more to working engineers than a benchmark percentage, but they're still self-reported and anecdotal — a single internal example, not a documented test set.

Three Speeds, One Model

SWE-2 ships in three effort tiers — medium, high, and max — letting Devin users trade cost against thoroughness depending on the task. Medium is positioned as the cheap, fast option; Cognition has not published comparable cost or benchmark breakdowns for the high and max tiers in what it's released so far. The question is whether an independent benchmark, run by someone other than Cognition, will confirm the 64% cost gap against Anthropic's model holds up outside Cognition's own test conditions. Until that happens, the number is a vendor's claim about its own product.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

unknown
Tech TimesCognition SWE-2 Beats Frontier Coding AI at 64% Lower Cost Using Single-Run RL Training - Tech Times