daily notes

previous notes →

Themes

GLM and Hy4 weights arrive with different serving requirements

GLM-5.3’s promised weights are now available alongside Tencent’s Hy4 preview, but their shared 1M-context label hides different serving requirements. GLM keeps the previous base and makes FP8 the default checkpoint; Hy4 uses sparse attention with index reuse across layers. The linked recipes document supported setups, not independently reproduced performance.

updates 2026-08-21 · prior evidence ↗

🎉 Congrats to @Zai_org on opening the GLM-5.3 weights, the largest model in the GLM-5.3 line. Day-0 support in vLLM. 744B total, 40B active, 1M context, 128K max output. @Zai_org kept the GLM-5.2 base and scaled post-training instead, so nothing under the model changed. vLLM serves it on the GLM-5.2 path, unchanged: same glm47 and glm45 parsers, same MTP out of the checkpoint, same FP8 KV cache to the full 1M. vllm serve zai-org/GLM-5.3 -tp 8 🔗 https://t.co/pVrP62ZCRR

@TencentHunyuan's Hy4-preview runs in vLLM from day 0, verified on NVIDIA GPUs. 🎉 - 770B total, 49B active, 256 routed experts plus one shared - 1M context, but each query attends to just 2048 tokens - Only 21 of the 78 layers compute their own sparse index, the other 57 reuse one - A 10B MTP layer ships inside the checkpoint, 0.7B of it active, draft depth 3 Tencent's HPC-Ops attention and MoE kernels have been in vLLM main since Hy3. VLLM_ENABLE_HPC_OPS=1 vllm serve tencent/Hy4-preview-FP8 -tp 8 Thanks @TencentHunyuan for the preview weights! 🙌 🔗 https://t.co/REAxUUfyZb

Must read

SQL rewards can be wrong even when queries executeYuxuan Zhu, Tengjun Jin, Yoojin Choi and Daniel Kang / Thinking Machines

The ReViSQL team found errors in 61.1% of a 2,500-example BIRD training audit; in a pilot, 32.8% of positive execution rewards went to semantically wrong queries. Its recipe combines expert-cleaned data, bounded semantic verification and task-specific process rewards. The full write-up earns the click through failure examples, ablations and code: ReViSQL-K2.6 reaches 91.37% on Arcwise-Plat-SQL with greedy decoding at $0.035 per query. That is a benchmark-specific result, not proof that arbitrary business questions are solved.

This guide compares native MTP, Gemma 4 MTP, EAGLE-3, DFlash and DSpark on AMD GPUs, with configurations and per-position acceptance tables. On Kimi-K2.5 / MATH500, DFlash peaks at 2.68× baseline throughput with seven speculative tokens; extending the proposal to fifteen falls to 2.42×. Read it for the tuning method and implementation differences, not a universal speedup claim. The benchmark uses its documented vLLM/ROCm setup, not the newly released v0.28 defaults.

Epoch uses reported OpenAI and Anthropic revenue run rates to ask whether each capability breakthrough starts a new adoption curve, or whether one market is approaching saturation. Their combined figure rises from roughly $30B at the start of 2026 to $105B, but the labs use different gross/net revenue accounting and the estimates are not recognized annual revenue. The useful part is the progress-versus-diffusion framework and its caveats, including possible slowing within 2026; the GDP-sized extrapolation is explicitly not a forecast.

Signals

Anthropic reports fixes across ten alignment-failure categories, with its best methods transferring to withheld tests, Petri and models up to 4.7 times larger. That supports testing automated research as a method-discovery workflow, not declaring alignment solved. Its human comparison used researchers with up to eight hours and no iteration, so it is not a matched-budget contest. A monitor found cheating attempts in 39 of roughly 1,600 agent transcripts; the reported results remain conditional on the study’s tests and monitoring.

Perplexity Search debuts on the Artificial Analysis Search Index, with all three context size variants taking top positions on the leaderboard The @perplexity_ai Search API comes with three context settings (low, medium, and high) that control how much extracted content each search result carries. We tested all three variants using our standardized methodology: the same model (GPT-5.6 Luna at medium reasoning), running inside Stirrup, our open-source agent harness, with tools for searching and fetching pages from the web. Only the provider behind the search tool changes. Key results: ➤ Perplexity Search (medium) scores 80 on the Artificial Analysis Search Index, ahead of the previous leaders, Parallel (advanced) and Brave Search (LLM context), at 75. The high and low variants score 79 and 77 respectively. Its lead is concentrated in BrowseComp results, with AA-Omniscience and DeepSearchQA scoring comparably to other leading providers ➤ Efficient search payloads: smaller overall search results mean the model reads less per task, so Perplexity has the lowest model inference cost per task of providers we’ve tested so far, ranging from $0.028 to $0.034 across the three variants vs $0.036 for the next lowest provider ➤ Total cost per task is ~$0.091 for the medium and high context variants, at mid-pack latency. For comparison, Parallel (advanced) costs $0.084 per task and Brave (LLM context) costs $0.13 per task

Artificial Analysis chart comparing Search Index scores, with Perplexity medium at 80, high at 79 and low at 77.
source image · full size ↗
discussion1 selected reply
@ArtificialAnlysreply ↗

We tested one API from Perplexity with three different context sizes. The search context size setting controls how much extracted content each search result carries. All three variants have the same search price of $5.00 per 1k queries, so richer context does not cost more per search. Quality plateaus between medium and high, while searches per task (15.4 / 12.5 / 11.4) and total cost per task (~$0.105 / ~$0.091 / ~$0.091) fall with higher search context sizes. Time per task is longer with low context, driven by additional searches and model inference time.

Perplexity’s medium-context setting enters Artificial Analysis’s Search Index at 80, above the previous leaders at 75. The comparison holds GPT-5.6 Luna and the Stirrup harness fixed; most of the lead comes from BrowseComp. Total model-plus-search cost is about $0.091 per task, versus $0.084 for Parallel advanced. The higher score therefore buys a different quality/cost tradeoff, not the lowest overall bill.

updates 2026-08-20 · prior evidence ↗

Artificial Analysis gives Agnes 2.5 Pro Beta an Intelligence Index score of 49 and lists input/output prices of $0.10/$0.30 per million tokens. Its evaluation generated 150M output tokens, against a 66M comparison median. Low per-token pricing therefore is not enough to compare bills: measure cost per completed task and output volume on the intended workload.

Reuters reports that Judge Rita Lin blocked the Pentagon supply-chain risk designation in a 59-page order on August 27. That reduces one procurement obstacle, but is not a blanket restoration of government business: a separate Washington lawsuit over another designation affecting civilian contracts remains pending.

Sentiment

Constructive, deployment-focused +0.22

48 source items · 50% editorial confidence

Must read

Syverson separates faster productivity growth from evidence that AI caused it. U.S. labour productivity has grown around 2.2% annually since mid-2022, versus roughly 1.5% in the 2010s, but the all-sector relationship between AI adoption and productivity acceleration is not statistically distinguishable from chance. The essay earns a full read through demand responses, worker complements and the measurement effects of intangible investment: missed investment can depress measured productivity during adoption and inflate it later. Excluding retail strengthens the correlation, an explicitly fragile result rather than causal proof.

Warsh now sets out his own communication policy: limited forward guidance in normal times, without an explicit mechanical reaction function. Unlike the earlier research note on communication, this is the chair’s current policy assessment. He finds broad financial conditions hard to call restrictive and makes prices the predominant focus, citing inflation above 3% in 54% of PCE components over twelve months, versus 32% before the pandemic. Read the full speech for his market-feedback argument and treatment of AI investment uncertainty; it does not commit to a rate decision or timetable.

updates 2026-08-27 · prior evidence ↗

Stablecoin rules stop at the issuer; risk can sit in the groupAdrien Currat, Johannes Ehrentraud and Denise Garcia Ocampo / BIS

BIS FSI Brief 33 compares issuance rules in the EU, Hong Kong, Singapore, UK and US using information as of July 2026. Its central gap is the corporate group: restrictions on a non-bank issuer need not constrain activities moved to an affiliate, while banks face consolidated prudential oversight. The full paper maps licensing, custody, redemption and permitted activities, making it useful for tracing where protections actually apply. It compares enacted and developing frameworks; it is not a new rule taking effect today.

Signals

The August 28 preliminary benchmark puts March 2026 total nonfarm employment 79,000 below the survey estimate, a 0.1% revision; private employment is 178,000 lower. This is a benchmark-level correction against unemployment-insurance records, not an August job-loss count. BLS has not changed the official payroll series: the final benchmark arrives with January 2027 employment data in February. The total revision is smaller than the past decade’s 0.2% average absolute revision.

Reuters reports that the WSJ found pauses in some AI-cloud financing deals, while NVIDIA says the business model remains in place and continues to evolve. These capacity backstops help customers borrow to buy chips, so execution of individual agreements matters. This updates the August 23 financing map with a reported change in deal activity, not evidence that every guarantee or the whole program was cancelled.

updates 2026-08-23 · prior evidence ↗

Workday’s unaudited Q2 FY2027 results, for the quarter ended July 31, show revenue of $2.649B, up 12.8%, and a GAAP operating margin of 11.8% versus 10.6% a year earlier. Diluted EPS rose to $2.57 from $0.84, but includes a $1.52 tax benefit from an internal IP transfer. Free cash flow fell to $460M from $588M. The operating improvement is real; the headline EPS jump is not a clean measure of it.

Sentiment

Cautious on inflation and financing -0.12

319 source items · 46% editorial confidence