daily notes

previous notes →

Themes

Qwen put Max-scale weights on the clock

Qwen now has a dated open-weight path and an independent performance read showing agentic gains, but it does not lead Kimi K3 on score or cost and the weights are not yet deployable. Artificial Analysis scored Qwen3.8 Max at 53, seven points above Qwen3.7 Max, but Kimi K3 remained four points ahead at roughly half the per-task cost; deployable weights and vLLM day-zero support remain promises until next week, with observed gains in agentic work rather than scientific reasoning.

Qwen3.8-Max is live: $2/$6 per MTok, flat across the full 1M context. One gotcha: reasoning_effort defaults to xhigh, and thinking tokens bill as output. We hit the endpoint in Apidog: it parses the SSE stream and shows the reasoning deltas separately in the timeline, so you can see what xhigh burns before the answer even starts. Then decide if medium is fine. Weights next week.

vLLM will support Qwen3.8-Max's open-source as usual on day-0, stay tuned🥰

Alibaba's Qwen3.8 Max makes real progress on agentic work tasks, scoring 53 on the Artificial Analysis Intelligence Index at $1.76 per task, but open weights leader Kimi K3 remains 4 points ahead at half the cost per task ($0.86) @Alibaba_Qwen has released Qwen3.8 Max, which Alibaba states is a 2.4T total parameter MoE activating 95B parameters per forward pass. The planned weights release would end Alibaba's pattern, in place since Qwen2.5 Max (January 2025), of shipping Max models as closed weights, and at 2.4T parameters would be ~6x larger than Alibaba's largest open weights release to date (Qwen3.5 397B) and second in size only to Kimi K3 (2.8T) Key takeaways: ➤ Qwen3.8 Max scores 53 on the Artificial Analysis Intelligence Index, up 7 points from Qwen3.7 Max (46). It is in line with Claude Sonnet 5 (max, 53) and sits second among Chinese labs, ahead of GLM-5.2 (max, 51) but behind Kimi K3 (max, 57) ➤ Qwen3.8 Max scores 1599 Elo on GDPval-AA, a 329 Elo gain over Qwen3.7 Max and its largest step forward. This places it effectively tied with Claude Sonnet 5 (1600) and ahead of GLM-5.2 (max, 1510), though behind Kimi K3 (1687) ➤ On AA-Briefcase, our benchmark for long-horizon knowledge work, Qwen3.8 Max scores 1430 Elo, ahead of Claude Sonnet 5 (max, 1385) and GLM-5.2 (max, 1253). Only Claude Opus 5 (max, 1721), Claude Fable 5 (1574), Kimi K3 (max, 1542) and GPT-5.6 Sol (max, 1504) score higher ➤ Remaining gains over Qwen3.7 Max are in agentic coding, with Terminal-Bench v2.1 up 6 points. Scientific reasoning is close to flat (HLE +2 points, CritPt +5 points, GPQA unchanged), while SciCode (-4 points), AA-LCR (-4 points) and AA-Omniscience (-11 points, driven by a hallucination rate rising 23% to 40%) regress ➤ Qwen3.8 Max costs $1.76 per Intelligence Index task, equivalent to Claude Sonnet 5 ($1.72). This is ~2x Kimi K3 (max, $0.86) and ~3x GLM-5.2 (max, $0.59), with only Claude Opus 5 (max, $2.34) and Claude Fable 5 ($3.15) costing more among leading models. Cost is driven partly by rising token usage: Qwen3.8 Max used 150M output tokens to run the Intelligence Index, up 50% from Qwen3.7 Max (100M), though half of Claude Sonnet 5 (300M) ➤ The 𝜏³-Bench Banking result (40%) appears out of distribution. It is a 29 point gain over Qwen3.7 Max, a large jump, and places Qwen3.8 Max ahead of models that outscore it on other agentic evaluations Key model details: ➤ Size: 2.4T total parameters, ~95B active per forward pass (MoE) ➤ Context window: 1M tokens ➤ Multimodal: text, image and video input with text output ➤ Pricing: $2.00/$6.00 per 1M input/output tokens on the @alibaba_cloud first-party API, with a $0.25 per 1M cache hit price ➤ Availability: Alibaba Cloud first-party API. Alibaba states the weights will be released next week

Artificial Analysis benchmark comparison chart for Qwen3.8 Max
source image · full size ↗
discussion1 selected reply
@ArtificialAnlysreply ↗

On AA-Briefcase, Qwen3.8 Max scores 1430 Elo overall, driven by analytical quality rather than presentation: its Analytical Quality Elo (1595) is 255 points above its Presentation Elo (1340). It passes 48% of rubric checks, ahead of Claude Sonnet 5 (max, 42%) and GPT-5.6 Sol (max, 42%) but behind Kimi K3 (max, 51%), and completes all checks on 3 of 89 tasks

Must read

OpenAI published ten manuscripts spanning geometry, cryptography, complexity, and related fields, alongside formal Lean certificates and reasoning walkthroughs. The artifacts make the first-party claim testable; the reported run cost was roughly $2,000 at GPT-5.6 Sol API rates.

The full-duplex system keeps media on a dedicated fast path while reasoning and tools run asynchronously. Its WARP protocol reduced voice-session startup from six network round trips to one, exposing CPU-side streaming and session management as scaling constraints.

Signals

Epoch's updated leaderboard put Claude Fable 5 at a 64% solve rate and GPT-5.6 Sol at 20% on whole-program reimplementation with held-out tests. Epoch flags possible pretraining contamination, so the gap is informative without being a clean measure of autonomous software engineering.

Sentiment

artifacts and systems, weights pending +0.42

25 posts · 90% confidence

Themes

No themes have been published for this stream yet.

Must read

No must-reads have been published yet.

Signals

Seven producers will add 188,000 bpd in September by restoring part of the April 2023 voluntary cuts. Monthly review continues, and the group still requires full compensation for overproduction since January 2024.

Julie Van de Kamp reports that the overall SONAR Truckload Rejection Index (STRI/OTRI) sits at 14.36%—significantly above its 6-month baseline average of 10.9%. Specialized equipment leads the tightening trend, with flatbed tender rejections at 23.5% and refrigerated (reefer) rejections at 19.46%. Because elevated rejection levels have persisted over months rather than temporary seasonal spikes, Julie emphasizes that market capacity is steadily shrinking rather than quickly re-balancing.

SONAR's composite truckload rejection index is 14.36% versus a 10.9% six-month baseline; flatbed and reefer rejections reached 23.5% and 19.46%. Persistent elevation points to capacity tightening even before a broad demand rebound.

Manufacturing expansion accelerated in JulyInstitute for Supply Management

The final July PMI rose 2.3 points to 55.6, its highest reading since May 2022 and seventh month of expansion. Production jumped to 58.5 and employment returned to growth at 52.8, while prices remained elevated at 71.1.

The final July SLOOS, covering Q2 and 56 domestic banks plus 18 foreign-bank branches, found unchanged C&I standards with stronger large-firm demand and easier CRE standards. Credit-card standards tightened, and standards for every surveyed NDFI loan type remained at the tight end of its post-2011 range.

GDPNow raised its Q3 growth estimate to 6.2%Federal Reserve Bank of Atlanta

The August 3 preliminary nowcast rose from 5.0% on July 30 to a 6.2% seasonally adjusted annual rate. Its inputs now imply 4.6% consumption growth and 17.9% private-investment growth, a larger handoff from the 1.5% advance Q2 estimate.

Sentiment

growth accelerating, credit selective +0.24

285 posts · 90% confidence