We independently benchmarked Devin Fusion for its release today - this is the first time a multi-model coding agent has been included on the Artificial Analysis Coding Agent Index, and it effectively retains Claude Fable 5.1 and GPT-6 Astra performance while reducing…
We independently benchmarked Devin Fusion for its release today - this is the first time a multi-model coding agent has been included on the Artificial Analysis Coding Agent Index, and it effectively retains Claude Fable 5.1 and GPT-6 Astra performance while reducing costs Devin Fusion runs a frontier lead model with a cost-efficient sidekick. We tested configurations from Cognition combining frontier models from Anthropic and OpenAI with their new SWE-2 (medium) as a sidekick model. Configured with Claude Fable 5.1 (xhigh) + SWE-2 (medium), Devin Fusion scores 62 on the Coding Agent Index v1.5, while with GPT-6 Astra (xhigh) + SWE-2 (medium) it scores 59. The Fable configuration has the higher score, while the Astra configuration is 43% less expensive and completes tasks 31% faster. Congratulations to @cognition on the release! See below for our results and analysis 🧵
discussion1 selected reply
Artificial Analysis reports Coding Agent Index scores of 62 for a Claude Fable configuration and 59 for a GPT-6 Astra configuration, with the latter 43% less expensive and 31% faster in its test. The comparison is specific to the tested configurations and benchmark.
Devin Fusion performs well for cost efficiency and performance, and currently sits on the Pareto frontier for Coding Agent Index score vs. Cost per Task in both configurations we tested Devin Fusion CLI with Claude Fable 5.1 (xhigh) + SWE-2 (medium) scores 61.7 on the Artificial Analysis Coding Agent Index v1.5. This is almost tied with Claude Fable 5.1 (max, with fallback) in Claude Code at 62.2 despite the lower effort, and costs 36% less at $7.9 per task for the Fusion configuration vs. $12.4 for Claude Code. Speed is also essentially flat, with time per task of 35.8 vs. 34.8 minutes. This pattern holds across the underlying evaluations: Fusion scores 63.1 vs. 64.3 on DeepSWE 1.1, 65.9 vs. 64.8 on SWE-Atlas QnA, and 56.1 vs. 57.6 on Terminal-Bench 4.0.