daily notes

previous notes →

Signals

Announcing Artificial Analysis Capability Indices v1.1, updated with stronger domain tuning, combining slices of core Intelligence Index v4.3 evaluations alongside specialized evaluations The Capability Indices map tasks from O*NET occupations to benchmarks that represent them, weighting each benchmark by how often its capability appears across the work. We cover six indices: Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical, Engineering, and Economics Key changes: ➤ Finance & Accounting: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Finance), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context Reasoning, and AA-Briefcase added to Agentic Knowledge Work ➤ Strategy & Ops: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Operations), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context Reasoning, and AA-Briefcase added to Agentic Knowledge Work ➤ Legal: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Operations and Support), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context Reasoning, and AA-Briefcase added to Agentic Knowledge Work ➤ Healthcare & Medical: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Operations and Support), Long-Context Reasoning added as a new capability sourced from MLCR-AA, Agentic Customer Interaction removed from capabilities, and AA-Briefcase added to Agentic Knowledge Work ➤ Engineering: Terminal-Bench updated to v4.0 in Agentic Terminal Use, GPQA Diamond removed from Reasoning, and AA-Briefcase added to Agentic Knowledge Work ➤ Economics: AA-Briefcase added to Agentic Knowledge Work Key results: ➤ Claude Fable 5.1 (max) leads in all six indices, followed by GPT-6 Astra (max) in Finance & Accounting, Strategy & Ops, Legal, and Engineering ➤ Open weights models are competitive across the Capability Indices. Kimi K3 (max) leads in Finance & Accounting (#8), Legal (#9), and Economics (#7), while DeepSeek V4.1 Flash (max) leads in Strategy & Ops (#7), and GLM-5.3 (max) leads in Healthcare & Medical (#6) and Engineering (#7)

Artificial Analysis says its new indices map O*NET occupations to specialized benchmarks across six domains and reports Claude Fable 5.1 leading all six. The methodology and ranking are those described by the evaluator, not an independent audit here.

OpenAI’s GPT-Live-1 debuts at #1 on the Artificial Analysis Speech to Speech Index at a score of 81.5 using Astra as its delegated backend model, ahead of Grok Voice Think Fast 2.0 GPT-Live-1 is @OpenAI's new full duplex Speech to Speech model that can delegate reasoning and tool use to a backend text model while continuing the conversation. Developers stream audio in and receive speech back through the API, with the backend text model configured separately. We evaluated two backend configurations: Astra at medium reasoning effort and Sol at low reasoning effort. Key takeaways: ➤ Speech to Speech Index: GPT-Live-1 (Astra, medium) achieves 81.5, ranking #1, while GPT-Live-1 (Sol, low) scores 80.1, ranking #3. Grok Voice Think Fast 2.0 High sits between them at 81.3 ➤ Speech Agent Arena: GPT-Live-1 (Sol, low) ranks #3 in preference at 1,053 Elo with 90.9% Task Success Rate, while GPT-Live-1 (Astra, medium) ranks #4 at 1,048 Elo with 87.4% task success. Gemini 3.1 Flash Live Minimal leads preference at 1,096 Elo, while Grok Voice Think Fast 2.0 High leads task success at 94.6% ➤ Tau Voice: GPT-Live-1 (Astra, medium) and GPT-Live-1 (Sol, low) take the top two spots on our agentic-performance benchmark at 67.9% and 59.3%, respectively, ahead of Grok Voice Think Fast 2.0 High at 56.5%. ➤ Big Bench Audio: GPT-Live-1 (Astra, medium) scores 90.1% and GPT-Live-1 (Sol, low) scores 89.0% on audio reasoning, behind Grok Voice Think Fast 2.0 High at 97.2% and Qwen Audio 3.0 Realtime Plus at 99.2% ➤ Speed: Average Time to First Audio on Big Bench Audio is 1.34 seconds for GPT-Live-1 (Astra, medium) and 1.24 seconds for GPT-Live-1 (Sol, low), compared with 0.70 seconds for Grok Voice Think Fast 2.0 High ➤ Cost: GPT-Live-1 (Astra, medium) costs $5.83 per hour of input audio, compared with $4.47 for GPT-Live-1 (Sol, low), including delegated backend model usage, versus $4.80 for Grok Voice Think Fast 2.0 High on our fixed Big Bench Audio pricing subset See below for more detail ⬇️

Artificial Analysis reports GPT-Live-1 with Astra at 81.5 on its Speech to Speech Index and gives separate task-success, latency, and cost comparisons. The reported ranking depends on the evaluator's selected tests and backend configurations.

Sentiment

Capability evidence with deployment constraints +0.18

66 source items · 55% editorial confidence

Must read

Bank executive compensation and risk-takingGaston Gelos, Bertrand Rime, and Kevin Tracol

The article links cross-jurisdiction compensation differences to risk-taking and finds that longer deferral periods are associated with lower bank risk.

Sentiment

Rates, credit, and infrastructure risk -0.08

426 source items · 55% editorial confidence