Dario Amodei's essay proposes pacing AI development through embedded evaluators, democratic coordination, and regulatory frameworks balancing innovation with safety.
daily notes
previous notes →Must read
Signals
Google has released Gemini 3.8 Live, its new Speech to Speech model, with the Extended Thinking (High) variant debuting at #1 on the Artificial Analysis Speech to Speech Index at 82.6, and #1 on our Tau Voice benchmark implementation at 68.6% Gemini 3.8 Live is…
Google has released Gemini 3.8 Live, its new Speech to Speech model, with the Extended Thinking (High) variant debuting at #1 on the Artificial Analysis Speech to Speech Index at 82.6, and #1 on our Tau Voice benchmark implementation at 68.6% Gemini 3.8 Live is @GoogleDeepMind's successor to Gemini 3.1 Flash Live, a Speech to Speech model that executes tools and API calls in the background while continuing the conversation. It comes in two variants: the standard Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, which supports configurable reasoning effort. We evaluated the standard model and the Extended Thinking variant at High reasoning effort through the Gemini Live API. Key takeaways: ➤ Speech to Speech Index: Gemini 3.8 Live Extended Thinking (High) debuts at #1 at 82.6, ahead of GPT-Live-1 (Astra, medium) at 81.5, Grok Voice Think Fast 2.0 High at 81.3 and GPT-Live-1 (Sol, low) at 80.1. The standard Gemini 3.8 Live debuts at #5 at 76.0, with both variants up on Gemini 3.1 Flash Live High at 71.5 (+11.1 and +4.5 points). The Index averages Speech Reasoning (Big Bench Audio), Agentic Performance (Tau Voice), Arena Preference and Arena Task Success Rate ➤ Speech Agent Arena: Gemini 3.8 Live ranks #2 in preference at Elo 1083, behind Gemini 3.1 Flash Live (1096) and ahead of GPT-Live-1 (Sol, low) at 1053, and #2 on Task Success Rate at 93.2%, behind Grok Voice Think Fast 2.0 High at 94.6%. Gemini 3.8 Live Extended Thinking (High) trails at Elo 990 with 89.1% task success ➤ Tau Voice: Gemini 3.8 Live Extended Thinking (High) takes the top spot on our Tau Voice benchmark implementation at 68.6%, ahead of GPT-Live-1 (Astra, medium) at 67.9%, GPT-Live-1 (Sol, low) at 59.3% and Grok Voice Think Fast 2.0 High at 56.5% - up from 37.7% for Gemini 3.1 Flash Live High. The standard Gemini 3.8 Live scores 30.1% ➤ Big Bench Audio: Gemini 3.8 Live Extended Thinking (High) scores 97.7% on audio reasoning, ahead of Grok Voice Think Fast 2.0 High at 97.2% and behind Qwen Audio 3.0 Realtime Plus at 99.2%. The standard Gemini 3.8 Live scores 91.7% ➤ Speed: Average Time to First Audio on Big Bench Audio is 1.18 seconds for Gemini 3.8 Live and 1.35 seconds for Extended Thinking (High), both well ahead of Gemini 3.1 Flash Live High (2.99s) and in line with GPT-Live-1 (Sol, low) at 1.24s and GPT-Live-1 (Astra, medium) at 1.34s, though behind Grok Voice Think Fast 2.0 High at 0.70s ➤ Cost: Gemini 3.8 Live costs $0.84 per hour of input audio, the cheapest model in the Index and roughly half the $1.75 of Gemini 3.1 Flash Live High. Extended Thinking (High) costs $3.50 per hour - cheaper than GPT-Live-1 (Sol, low) at $4.47, Grok Voice Think Fast 2.0 High at $4.80 and GPT-Live-1 (Astra, medium) at $5.83, and ~3.1x cheaper than GPT-Realtime-2.1 High at $10.75 See below for more detail ⬇️
Artificial Analysis reports Gemini 3.8 Live Extended Thinking at 82.6 on its Speech to Speech Index and 68.6% on its Tau Voice implementation, while listing speed and hourly-cost comparisons. The figures are benchmark-specific and include separate standard and extended variants.
Introducing Retrieve-for-Train: a framework that accelerates complex AI search by replacing heavy autoregressive inference with a lightweight diffusion model. This allows for instant, expert-level search slates at a fraction of the cost. Learn more: https://t.co/cQP0Lmg5kj…
Introducing Retrieve-for-Train: a framework that accelerates complex AI search by replacing heavy autoregressive inference with a lightweight diffusion model. This allows for instant, expert-level search slates at a fraction of the cost. Learn more: https://t.co/cQP0Lmg5kj https://t.co/xQsdpujqIF
Google Research says Retrieve-for-Train replaces heavy autoregressive inference with a lightweight diffusion model to accelerate complex search. The post states a cost advantage but gives no quantitative result, so the implementation claim remains unquantified here.
Sentiment
87 source items · 55% editorial confidence
Must read
The SemiAnalysis study uses its datacenter model to distinguish 20 GW under restriction from roughly 1,525 MW actually delayed, while noting future opposition risk.
Sentiment
437 source items · 55% editorial confidence