Epoch separates Huawei's current AI-compute constraint into chip performance, production volume and HBM supply. Even ten times China's projected usable HBM would put Huawei at only about 11% of Nvidia's 2028 output; by then the modeled bottleneck shifts toward compute per gigabyte, while Huawei's first LogicFolding Ascend chip is not scheduled until 2030.
daily notes
previous notes →Must read
Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming Intelligence Index…
Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming Intelligence Index v4.2 changelog: + AA-Briefcase, our agentic knowledge work evaluation with a private test set + @HelloSurgeAI's GDP.pdf, long context document reasoning across 4,592 PDF pages - GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated … plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January. We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users. Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned! Intelligence Index v4.2 changes in detail: ➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work. ➤ Adding GDP.pdf: Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied. ➤ Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5. ➤ Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure. Key results: ➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z AI, and Google ➤ Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z AI occupy the updated Cost per Task Pareto frontier ➤ GPT-6 Astra dominates the output token Pareto frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)
Intelligence Index v4.2 replaces saturated GPQA Diamond with harder work: AA-Briefcase tests multi-week agentic projects, GDP.pdf spans 4,592 document pages, and private held-out sets now carry 40% of the index, double v4.1. The full note also records grading and sampling repairs, making the rank changes interpretable instead of a fresh leaderboard without lineage.
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment…
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact. For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways. Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported in https://t.co/9aiRxk2eUJ, https://t.co/ADjyzwSUGz, and https://t.co/SUV6jZ3Gaz. We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared. Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.
OpenAI says real-world agent misalignment now needs incident disclosure, not only research write-ups. It distinguishes model properties from deployment impact, cites its Hugging Face security response and earlier unintended internet use, and promises a framework in the coming weeks; that framework is not yet public.
Signals
A new REST endpoint returns historical repository star counts with timestamps but not individual stargazer identities. It replaces a common analytics use case lost when GitHub restricted person-level stargazer listings to admins and collaborators.
Sentiment
75 source items · 76% editorial confidence
Must read
Using 2022 Survey of Consumer Finances responses for tax year 2021 and TAXSIM, the authors estimate effective federal rates from -12.4% in the lowest income decile to 25.0% in the highest. The decomposition is the useful part: wage taxation drives most of that slope, while investment and other nonwage rates stay much flatter across income groups.
The 10-year Treasury yield reached 4.82%, its highest since 2023, while the spread on triple-C-or-lower debt widened to 10.53 percentage points from 8.08 a year earlier. Default actions are up 9% year to date to $40.1 billion and recoveries average 29%, showing a sharply bifurcated credit market rather than broad stress.
The share of US adults saying they were doing at least okay financially slipped only from 75% in 2019 to 73% in 2025, while Michigan consumer sentiment fell from a 2019 average of 95.9 to 71.7 in January 2025. The gap reflects different questions: present financial position versus change relative to a year earlier.
Signals
The Financial Times reports that Anthropic is close to awarding Morgan Stanley and Goldman Sachs top roles in an offering discussed at roughly a $2 trillion valuation. The bank selection would move the prospective listing from general intent toward transaction preparation, but no IPO or valuation is complete.
CENTCOM says it struck three IRGC-linked crude carriers on September 5 after ballistic missiles targeted two US warships. The US reported no personnel or vessel damage; two tankers were disabled and one was destroyed after crews abandoned them, adding a direct escalation risk to oil shipping.
Sentiment
212 source items · 79% editorial confidence