daily notes

previous notes →

Signals

We independently benchmarked Devin Fusion for its release today - this is the first time a multi-model coding agent has been included on the Artificial Analysis Coding Agent Index, and it effectively retains Claude Fable 5.1 and GPT-6 Astra performance while reducing costs Devin Fusion runs a frontier lead model with a cost-efficient sidekick. We tested configurations from Cognition combining frontier models from Anthropic and OpenAI with their new SWE-2 (medium) as a sidekick model. Configured with Claude Fable 5.1 (xhigh) + SWE-2 (medium), Devin Fusion scores 62 on the Coding Agent Index v1.5, while with GPT-6 Astra (xhigh) + SWE-2 (medium) it scores 59. The Fable configuration has the higher score, while the Astra configuration is 43% less expensive and completes tasks 31% faster. Congratulations to @cognition on the release! See below for our results and analysis 🧵

discussion1 selected reply
@ArtificialAnlysreply ↗

Devin Fusion performs well for cost efficiency and performance, and currently sits on the Pareto frontier for Coding Agent Index score vs. Cost per Task in both configurations we tested Devin Fusion CLI with Claude Fable 5.1 (xhigh) + SWE-2 (medium) scores 61.7 on the Artificial Analysis Coding Agent Index v1.5. This is almost tied with Claude Fable 5.1 (max, with fallback) in Claude Code at 62.2 despite the lower effort, and costs 36% less at $7.9 per task for the Fusion configuration vs. $12.4 for Claude Code. Speed is also essentially flat, with time per task of 35.8 vs. 34.8 minutes. This pattern holds across the underlying evaluations: Fusion scores 63.1 vs. 64.3 on DeepSWE 1.1, 65.9 vs. 64.8 on SWE-Atlas QnA, and 56.1 vs. 57.6 on Terminal-Bench 4.0.

Artificial Analysis reports Coding Agent Index scores of 62 for a Claude Fable configuration and 59 for a GPT-6 Astra configuration, with the latter 43% less expensive and 31% faster in its test. The comparison is specific to the tested configurations and benchmark.

Hi Astra users. A reset and a quick update on quality issues that have been posted around. Working with some of you, we have found and fixed the following issues: - Some skills written for previous models were triggering too often or preventing the model from checking its work. - An opt-in context management experiment that could cause early stops or replies to older messages. We've disabled it. Our rough estimate is that 4-5k users were affected by this experiment. - We've also removed some badly configured engines that resulted in a measured quality degradation for a long tail of traffic flowing through them. We’ve also made some more minor improvements and things should feel significantly better across the board. More consistent follow-through, better tracking of your latest message, and better checks on the work as it’s going through the motions. The examples posted and all the users who worked directly with us were incredibly useful in helping fix things quickly. Always grateful for this incredible community. And of course, a reset is also landing by midnight today.

The post says an opt-in context experiment affected an estimated 4–5k users, some older skills triggered too often, and badly configured engines degraded a long-tail of traffic. It says those issues were disabled or removed; the estimate remains rough.

Sentiment

Capability evidence with deployment constraints +0.18

53 source items · 55% editorial confidence

Must read

Signals

Investors have all but concluded the Federal Reserve will raise interest rates next week for the first time in three years. The harder question is what comes after that. Because almost no one at the central bank thinks a quarter-point increase will do much on its own to bring inflation down, a decision to raise rates next week would reflect a judgment that interest rates have been in the wrong place. If that is the case, one increase won’t fix it. “If we get a hike next week, certainly we’ll get additional ones,” said Richard Clarida, a former Fed vice chair who is now at Pimco. https://t.co/QfGuNL8YVa

The post says investors broadly expect a quarter-point Federal Reserve increase and quotes Richard Clarida that additional hikes could follow because one move would not materially reduce inflation. It reports expectations and commentary, not a central-bank commitment.

Sentiment

Rates, credit, and infrastructure risk -0.08

378 source items · 55% editorial confidence