ARC Prize's September 3 evaluation advances the earlier Sol context-retention result: Astra reaches 62.7% on Semi-Private with the Standard harness at max effort, versus 99.9% with a Provider Adapter at high effort. The full report separates reasoning settings, costs and action efficiency, and exposes replays of the model's symbolic world-building. Read it before comparing headline scores: these are different system configurations, and ARC Prize explicitly rejects benchmark saturation as proof of AGI.
The newly released benchmark decomposes 123 financial-research tasks into 701 criteria covering retrieval, source verification and calculation. Across its evaluated configurations, 201 of 1,150 correct-answer instances lack complete supporting evidence, a 17.48% rate. Its task examples, scoring rules and open data are useful for testing research agents that produce plausible numbers with the wrong vintage, unit or source. The automated judge was checked on 50 instances; the results remain specific to this small, expert-built task set.
Anthropic's September 4 report and public repository document the first complete machine-checked formalization of Fermat's Last Theorem in Lean 4: 13 million generated lines, 29,500 intermediate theorems and 60,475 modules checked by Lean's kernel. The full artifact explains the Prove2Me multi-agent theorem graph, verification with independent kernels and the failed attempts retained in roughly 7% of non-boilerplate code. Read it for the workflow and proof graph, not as evidence that AI discovered Wiles's mathematical argument; Claude formalized an existing proof with occasional high-level human direction.
Meta's Muse Image debuts at #4 on the Artificial Analysis Image Editing Leaderboard and takes #5 in Text to Image, including a position on the Pareto frontier for quality vs price
Muse Image is the first image model from Meta Superintelligence Labs, launched in Meta AI in July…
Meta's Muse Image debuts at #4 on the Artificial Analysis Image Editing Leaderboard and takes #5 in Text to Image, including a position on the Pareto frontier for quality vs price
Muse Image is the first image model from Meta Superintelligence Labs, launched in Meta AI in July and now available to developers on the Meta Model API. Meta positions it as an agentic image model: it invokes search and coding tools, self-refines its own generations, and composes from multiple references.
In the Artificial Analysis Image Arena, Muse Image debuts at #4 in Image Editing, behind Microsoft's MAI-Image-2.6-Preview and OpenAI's GPT Image 2 (high) and narrowly behind Google's Nano Banana 2. In Text to Image it takes #5, behind only GPT Image 2 (high), MAI-Image-2.6-Preview, Reve 2.1, and Google's Nano Banana 2.
Muse Image is available in the Meta AI app, on https://t.co/RAZKpGp53L, in Instagram Stories, in WhatsApp in limited countries, and to developers via the Meta Model API and partner platforms fal, Runway, and OpenRouter.
Congratulations to @AIatMeta, @alexandr_wang, & @finkd on the release!
See below for our analysis and example outputs of Muse Image in the Artificial Analysis Image Arena 🧵
Where Muse Image is strongest: closest to the category frontier in Lighting and Reasoning for Text to Image, and at the frontier in Scene & Style Edits and Identity Preserving Edits for Image Editing
Our Image Arena taxonomy measures 9 capabilities and 10 use cases for Text to Image, and 7 editing actions and the same 10 use cases for Image Editing. Capabilities measure the individual skills that go into making an image. Editing actions measure what an edit asks the model to do.
Muse Image is closest to the Text to Image category frontier in our Lighting and Reasoning capabilities, and sits at the Image Editing category frontier in Scene & Style Edits and Identity Preserving Edits.
➤ Lighting tests how light behaves: direction, falloff, catchlights, and shadow that match the source.
➤ Reasoning measures whether the model resolves prompts that require inference rather than literal depiction.
➤ Scene & Style Edit covers relighting, restyling, and changes to the background, setting, or overall treatment.
➤ Identity Preserving Edit covers changes to people, faces, and characters that must keep identity intact.
Artificial Analysis places Muse Image fourth in image editing and fifth in text-to-image preferences, with a position on its quality-versus-price frontier. This is a new independent measurement of a model launched in July, not a fresh launch. The arena result supports testing it against your own image workload; it does not establish factual accuracy or reliable tool use.
This September 3 working paper links firms and banks across ten Asian economies over 2005–2021, then compares trade-weighted and bank-lending-weighted exposure abroad. Its main finding is that zombie-firm spillovers reach advanced-economy inflation and growth through intermediate-goods import prices, while cross-border bank exposure has no comparable macroeconomic effect in the study. The channel tests and domestic-versus-foreign-bank split make the full paper worth reading; this is historical evidence, not a measurement of today's zombie share.
Setser and Weilandt use company disclosures to distinguish cash tax payments, provisions for current tax and the location of reported profits. Their September 3 analysis contrasts technology companies booking foreign-sales profits in Ireland with pharmaceutical structures shifting profits from U.S. sales abroad. The filing links and treatment of legacy liabilities repay a full read: a cash payment settling old tax obligations is not evidence that current U.S. profits face the same tax burden. The policy conclusion is the authors' argument, not a finding of illegality.
The September 3 comparison of six jurisdictions shows how small-bank regimes combine eligibility limits with different capital, liquidity and reporting requirements. The useful detail is the pairing: removing a market-risk requirement calls for a trading-book limit, while liquidity exemptions need funding-risk safeguards. Read the jurisdiction tables before treating regulatory simplification as a uniform capital release. This is a comparative policy study, not a new rule applicable to every bank.
The preliminary August estimate adds 162,000 nonfarm jobs, while July is revised from -23,000 to +21,000 and June from +20,000 to +31,000, lifting those two months by 55,000 combined. Unemployment is unchanged at 4.1%, participation edges up to 61.6%, and annual wage growth slows to 3.1%. The rebound removes the prior outright payroll contraction, but one release does not establish a new hiring trend and August remains subject to revision.
Norges Bank Investment Management recommends cutting the government-bond share of the fund's fixed-income benchmark from 70% to 50%, splitting the benchmark evenly between government and other bonds. It would use market-value rather than GDP weights, restore inflation-linked sovereigns and expand securitized, government-related and corporate exposure. This is a pending recommendation to the Ministry, not an executed Treasury sale; the official submission does not quantify a U.S. disposal.
Where Muse Image is strongest: closest to the category frontier in Lighting and Reasoning for Text to Image, and at the frontier in Scene & Style Edits and Identity Preserving Edits for Image Editing Our Image Arena taxonomy measures 9 capabilities and 10 use cases for Text to Image, and 7 editing actions and the same 10 use cases for Image Editing. Capabilities measure the individual skills that go into making an image. Editing actions measure what an edit asks the model to do. Muse Image is closest to the Text to Image category frontier in our Lighting and Reasoning capabilities, and sits at the Image Editing category frontier in Scene & Style Edits and Identity Preserving Edits. ➤ Lighting tests how light behaves: direction, falloff, catchlights, and shadow that match the source. ➤ Reasoning measures whether the model resolves prompts that require inference rather than literal depiction. ➤ Scene & Style Edit covers relighting, restyling, and changes to the background, setting, or overall treatment. ➤ Identity Preserving Edit covers changes to people, faces, and characters that must keep identity intact.