daily notes

previous notes →

Must read

I used Jev to delegate the entire tool call, execution, and next-step loop away from the main LLM. The setup was simple: On one hand we have the native vanilla Harness . The LLM sees the tool catalog, picked a tool, got the result, reasoned again, picked the next tool, and repeated. The other harness gave the LLM exactly one tool: `delegate_to_jev`. This acts as the starting state. Jev then handled the closed-set routing: which real tool to call next, what typed arguments to use, when to continue, and when to stop. The LLM only framed the task and summarized the final evidence. We ran this on a deterministic tool calling benchmark with 12 tasks: multi-step workflows, safety checks, prerequisite chains, and state-changing actions. Results: Baseline harness: - 10 / 12 scenarios passed - $0.01966 cohort cost - 49,364 outer-LLM tokens LLM + Jev delegation: - 12 / 12 scenarios passed - $0.01249 cohort cost - 5,363 outer-LLM tokens That is ~36.5% lower cost and ~89.1% fewer outer-LLM tokens. That is crazyyyyy!! This means, we can use the LLM for reasoning and synthesis. And use Jev for fast, typed, repeated decisions inside the tool loop.

A first-party comparison gave the main model one delegation tool and let Jev route a closed tool set; it passed 12/12 tasks versus 10/12 for the baseline, while reported cost fell 36.5% and outer-model tokens fell 89.1%. The test is small and self-reported, but it adds measured deployment evidence to Jev's earlier launch.

updates 2026-09-17 · prior evidence ↗

Sentiment

Measured orchestration gains, thin broader evidence +0.16

15 source items · 48% editorial confidence

Must read

One way to think about the next phase for AI labs: growth, power, and margins, with monetized inference $ per gigawatt connecting the figures. Using $65B for Anthropic's reported July run rate and assuming 2-5 GW of capacity gives $13–32.5B per total GW, or $32.5-81.3B per inference GW at the same 40% allocation. Both measures imply roughly 1.3-3.3× OpenAI's historical monetization on a blended basis. Year-end capacity paths per public reporting for 2025/26/27: >Anthropic: 1.5 → 5 → 10 GW. >OpenAI: 1.9 → 5 → 10 GW. Annualized revenue = total GW × inference share × revenue per inference GW. >Growth. Assume 40% serves customers at $50B annually per inference GW. That gives $20B per total GW, within Anthropic's July sensitivity range. At that ratio, 5 → 10 GW implies $100B → $200B of annualized revenue per lab. Hence people circling that for YE26 guesses. They were at ~$105B as of mid-year w/ more capacity coming online. Total revenue can still grow while blended revenue per GW falls. >Power. Those paths require combined net additions of ~6.5 GW in 2026 and 10 GW in 2027. The latter would absorb 33–50% of illustrative annual market additions of 20–30 GW. Delivery timing matters as much as announced capacity. >Margins. Utilization, output per watt and realized pricing determine revenue per GW. Better chips and software can offset cheaper tokens. At an assumed annual compute cost of $20–25B per total GW, a 40% inference allocation requires $50–62.5B of annualized revenue per inference GW to cover all compute, before other operating costs. (For an illustrative latest-generation GPU fleet, assume 2 kW of total facility power per GPU, including supporting equipment and cooling: ~500,000 GPUs per GW. At 85% billable utilization, that's ~3.7B GPU-hours annually, implying an effective paid compute price of ~$5.4–6.7 per GPU-hour.) The visuals separate disclosed history from assumptions. These are implied annualized run rates, not full-year revenue forecasts. Open to other views on inference allocation and how monetization changes with scale.

Using newly reported Anthropic and OpenAI capacity paths, Mathew models 5 to 10 GW per lab and estimates that a 40% inference allocation needs $50-62.5B of annual revenue per inference GW just to cover $20-25B of compute cost per total GW. The inputs are scenario assumptions, not forecasts, but they update his earlier takeoff stress test with current lab scale.

updates 2026-09-02 · prior evidence ↗

Signals

The September 17 preliminary release put starts at a 1.275M seasonally adjusted annual rate, down 2.6% from revised July and 1.2% year over year; permits fell 2.7% month over month to 1.394M but rose 3.5% year over year, while completions fell 11.9% month over month. The estimates remain subject to revision.

$SPX is expected to report Y/Y earnings growth of 28.9% for Q3 2026, which is above the estimate of 26.7% on June 30. #earnings, #earningsinsight, https://t.co/joygkDuBUL https://t.co/Q80HHUFYcA

FactSet says Q3 2026 S&P 500 earnings growth is now expected at 28.9% year over year, up from 26.7% on June 30. It is a consensus forecast, not reported earnings, but the upward revision runs against the usual pre-quarter estimate cuts.

Sentiment

Earnings optimism against weaker housing +0.04

49 source items · 56% editorial confidence