Jeff Dean, Sanjay Ghemawat, Quoc Le and Oriol Vinyals launched a company to automate experimental loops across machine learning, science and engineering. The founding page establishes the mission; financing and partnership claims remain outside the note until first-party disclosure.
daily notes
previous notes →Themes
No themes have been published for this stream yet.
Must read
WeatherNext used nearly 5,000 historical storms to improve joint forecasts of cyclone track, intensity and wind structure, providing more than 24 hours of additional warning on average in 2023–24 tests. The model is open source, but the result remains a retrospective evaluation rather than a live operational record.
Signals
Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index at $1.14 per task, but open weights leader Kimi K3 remains 1 point ahead at 25% lower cost per task ($0.86) @Alibaba_Qwen has released Qwen3.8 Max, which Alibaba states is a 2.4T total parameter MoE…
Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index at $1.14 per task, but open weights leader Kimi K3 remains 1 point ahead at 25% lower cost per task ($0.86) @Alibaba_Qwen has released Qwen3.8 Max, which Alibaba states is a 2.4T total parameter MoE activating 95B parameters per forward pass. Alibaba has announced it plans to release the weights next week, a shift in strategy as it has typically kept its Max class of models proprietary. Once released, Qwen3.8 Max would be ~6x larger than Alibaba's largest open weights release to date (Qwen3.5 397B) and the second largest open weights model behind Kimi K3 (2.8T) Note: we earlier published results showing Qwen3.8 Max scoring 53 on the Artificial Analysis Intelligence Index. Those runs were affected by intermittent issues on the endpoint we were evaluating, and we have re-run all evaluations on Alibaba's public API endpoint Key takeaways: ➤ Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, up 10 points from Qwen3.7 Max (46). It is in line with Claude Opus 4.8 (max, 56) and sits second among labs from China, ahead of GLM-5.2 (max, 51) but behind Kimi K3 (max, 57) ➤ Qwen3.8 Max scores 1739 Elo on GDPval-AA, a 468 Elo gain over Qwen3.7 Max. This places it ahead of Kimi K3 (1685), effectively tied with Claude Fable 5 (1743) and GPT-5.6 Sol (max, 1730), and behind only Claude Opus 5 (max, 1852) ➤ Gains over Qwen3.7 Max span agentic evaluations, scientific reasoning and coding: Terminal-Bench v2.1 +6 points, CritPt +7 points, SciCode +4 points and HLE +3 points, with GPQA unchanged. AA-LCR (-2 points) and AA-Omniscience (-10 points, driven by a hallucination rate rising 23% to 40%) regress ➤ Qwen3.8 Max costs $1.14 per Intelligence Index task, more than double Qwen3.7 Max ($0.53), at ~1.3x Kimi K3 (max, $0.86) and ~2x GLM-5.2 (max, $0.57). Cost is driven in part by more turns on agentic evaluations, with GDPval-AA input tokens rising ~15x over Qwen3.7 Max, and output token usage up 45% to 145M ➤ The 𝜏³-Bench Banking result (42%) appears out of distribution. It is a 32 point gain over Qwen3.7 Max and places Qwen3.8 Max ahead of models that outscore it on other evaluations Key model details: ➤ Size: 2.4T total parameters, ~95B active per forward pass (MoE) ➤ Context window: 1M tokens ➤ Multimodal: text, image and video input with text output ➤ Pricing: $2.00/$6.00 per 1M input/output tokens on the @alibaba_cloud first-party API, with a $0.25 cache hit price. This is lower than Qwen3.7 Max across the board ($2.50/$7.50, with a $0.50 cache hit price) ➤ Availability: Alibaba Cloud first-party API. Alibaba states the weights will be released next week
The rerun corrected endpoint issues and raised Qwen3.8 Max from 53 to 56, narrowing Kimi K3's lead from four points to one while cutting measured cost per task from $1.76 to $1.14. The weights are still only promised for next week, and the measured hallucination rate still rose from 23% to 40%.
OpenAI made the updated Sol model the shared Instant and deep-reasoning model for Plus and Pro, while Free and Go users receive unlimited Luna text chats starting August 7. OpenAI reports 62% fewer factual errors for Luna and 68% fewer for Sol than GPT-5.5 Instant in internal evaluations; free-tier tools and media remain separately limited.
A weighted Ipsos KnowledgePanel survey of 1,106 employed U.S. adults found 20% had delegated at least one task to AI that they previously assigned to another person. The result is worker self-report collected July 10–19, not a measured change in employment or output.
Sentiment
35 posts · 88% confidence
Themes
Low layoffs now mask a slower exit from unemployment
The labor market's current weakness is in hiring and job-finding, not layoffs. Initial claims rose 1,000 to 199,000, while continuing claims rose 24,000 to 1.801 million; Cleveland and Richmond Fed analyses link this to low hiring and slower exits from unemployment.
Must read
No must-reads have been published yet.
Signals
KOSPI fell more than 4% and the Korea Exchange halted program selling for five minutes after the index's 14% clearing rally on July 30. The reversal makes one-session breadth a weaker basis for calling the forced AI unwind complete.
Japanese and U.S. companies plus Akita's prefectural and city governments are discussing one of Japan's largest AI data centers, with UAE investment and early-2030s operation targeted. The ¥2tn figure is a plan under negotiation, not committed financing.
June wholesale sales fell 3.0% to $794.1 billion after May's gain was revised to 3.5%, while inventories rose 0.2% to $944.7 billion. The inventory change was not statistically distinguishable from zero, and the inventory-to-sales ratio was 1.19 versus 1.30 a year earlier.
Renters expecting to move within three years fell from about 57% in 2014 to 37% in 2026, while their expectation of ever owning a home fell from 52% in 2015 to 35% in 2025. These are survey expectations and correlations, not realized moves or causal estimates.
Preliminary nonfarm productivity rose at a 1.4% annual rate in the second quarter as output increased 1.7% and hours 0.3%; unit labor costs rose 1.3%. The release is preliminary and includes revisions to first-quarter inputs.
Sentiment
308 posts · 89% confidence
