daily notes

previous notes →

Themes

Cursor’s acquisition split its model-provider support

Cursor’s acquisition by SpaceX has made model access a matter of provider routing, not a simple cutoff. OpenAI proposed ending direct access on November 12, creating a choice for users where API-key access and its Cursor extensions remain available, while Anthropic committed to increasing compute support for Claude, leaving the specific routing details and commercial terms unsettled.

We’re ending our partnership with Cursor following its acquisition by SpaceX. Under our proposal, Cursor’s direct access to our models would end on November 12. We know that the people most affected by this decision are the developers who rely on OpenAI models in Cursor. We care about their experience in this transition and we’re ready to go above and beyond to support them. https://t.co/OzuCTzUjfX

Cursor has been a trusted partner of Anthropic since Sonnet 3.5. We’ll continue to increase compute to support Claude models in Cursor and are excited for what comes next with them at SpaceX.

Must read

It is not inherently bad to publish research on the impact of AI that only refers to older models, but it requires a very careful discussion and has to be very clear to non-technical readers. If you show AI can do something, its fine. Generally, once AI has gained an ability, it does not regress in future generations. So papers that show that AI has crossed a threshold (such as "good as a human") or that seek to establish some sort of minimum impact or effect still hold up even if they used GPT-4. If the finding is that AI is bad or biased at something, you need to be much more careful. You cannot claim that because GPT-4 fails at something, that AI is bad at that task, because current or future models may do it. So what can you claim? These are just some examples: 1) You can just be explicit: "GPT-4 could not do X." This is a rare way to frame things because it is not really an interesting publishable paper in many cases. However, if you provide the information needed to reproduce your work, it can become a useful benchmark to measure progress. 2) You can measure trends: compare GPT-4 to GPT-5 to GPT-5.6 Sol or whatever (it is really important to include at least one reasoning model). You can then make better claims about abilities relating to this task. 3) You can make a strong, grounded argument that AI has a natural flaw or limitation that means it cannot do X, and then demonstrate evidence that this might be the case. 4) You can focus on a moderator or mediator: this sort of prompting or approach or social context or connection impacts the ability of GPT-4 to do X, and needs to be a concern for the future. 5) You can focus on humans: how do people react to AI? what are the failures and successes? what dangers or advantages or changes does it bring to us? Again, these aren't exhaustive, but generally negative capability claims have been much less durable than positive ones.

Mollick offers a durable rule for reading AI-impact research: positive capability thresholds often survive model turnover, while negative or bias findings need the exact model and date. He gives five better framings—reproducible benchmarks, trends across generations, grounded mechanism claims, moderators and human response—so the piece changes how to underwrite evidence rather than merely recapping one study.

Use LLMs to rank regulatory divergences, not decide themFernando Perez-Cruz, Kumar Rishabh and Ayush Uchil / BIS

The BIS turns an LLM into a supervisory triage layer: five independent GPT-5.5 Pro runs compare AT1 prospectuses with capital rules, then rank recurring clause-level divergences for human review. The procedure reduced 576 raw findings to 79 candidates recurring in at least four runs and recovered the Credit Suisse and Yes Bank wording differences after anonymisation. The authors also state the missing evidence clearly: expert-labelled tests are still needed to measure false negatives and decide how much full-document review remains necessary.

We are reseting usage for all paid users of Codex and ChatGPT Work. Please continue reading for an update on Codex usage limits. The team has been working around the clock, going through thousands of reports and shipping fixes. Depending on how you use Codex, you should see your usage go between 10% and 50% further than before. We really went with a fine comb, with many uncovered small things being longstanding and here is what we found and fixed: - Compaction. We were keeping old images during compaction, sometimes making the context large enough to trigger compaction again. After the fix, usage dropped around 10% for users making heavy use of images. Fixed. - Memory. Background memory workers could inherit Stop hooks and keep running when the hook wouldn’t let them finish. This affected fewer than 1% of users, with the long tail being pretty bad and we saw one example thread check whether it could stop 15,000 times. Fixed. - Goals. In some cases, a set /goal could finish and then keep going past the intended stop condition, or the model would keep retrying broken tools without stopping. We saw examples consume anywhere from 15% to 70% of a weekly allowance. Fixed. - Automations. Some custom schedules could run more frequently than configured. Fixed. - Subagents. Smaller models (e.g. Luna) sometimes picked more capable helpers without being explicitly asked. The same was true where the orchestrating model not running in /fast mode could request sub-agents to run /fast. Fixed. - Computer History. The older implementation could lead to repeatedly summarizing overlapping activity. For some cases we saw it consume up to one fifth of the weekly usage per week. Fixed. - Rolling task summaries. Ordinary turns were triggering extra background requests. These added about 1% to token usage. Small each time, but it adds up. We have disabled this. - MCP. Some tool results could be encoded twice. We also found tool instructions getting cut off and fetched again. Fixed. We’ve also made architectural changes to prevent these from regressing and our teams will get paged if it happens regardless. We are also working on showing you directly in the app where your usage goes so you don’t have to guess. Goes without saying that we’re resetting usage limits and I hope you enjoy a very nice Saturday!

This first-party operating note is worth the full read because the failure modes are unusually concrete: retained images could retrigger compaction, background memory workers could loop on Stop hooks, completed goals could consume 15–70% of a weekly allowance, computer-history summarisation could use one fifth, and MCP results could be encoded twice. OpenAI says the fixes should extend usage by 10–50% depending on workflow and reset paid-user allowances while adding regression paging.

Signals

A revised working paper compares the staggered geographic rollout of Google AI Overviews across the same Wikipedia articles and estimates a 5.45% reduction in English external-search referrals versus German and 4.82% versus French. The abstract supports a search-displacement effect, not a general estimate for all publisher traffic; the full paper was unavailable in this run, so treat the result as working-paper evidence.

Workplace AI use rose, but time savings remain self-reportedFRED Blog / Federal Reserve Bank of St. Louis

The Real-Time Population Survey puts workplace generative-AI use at 39.2% in Q2 2026, up from 28.2% in Q3 2024; assisted hours rose from 4.1% to 6.3% and reported time saved from 1.6% to 2.2% of work hours. The direction is useful adoption evidence, but the savings are respondents’ approximate self-reports rather than measured productivity.

GitHub plugin syncing is now available for ChatGPT workspaces, including Business and Enterprise. Admins can import a Codex or Claude marketplace from a public or private repo and share its plugins with their team. The workspace picks up changes daily, or sooner with Sync now, so there’s no need to upload a ZIP for every update. Existing installation and app-access controls still apply. If you use a personal account, you can already add plugin repos in the desktop app. Workspace admins can get started under Workspace settings → Plugins → Add → Import marketplace.

Business and Enterprise admins can now import a Codex or Claude plugin marketplace from a public or private GitHub repository, share it across the workspace and receive daily updates or trigger Sync now. Existing installation and app-access controls still apply, so the operational change is centralized distribution and updates rather than a permissions bypass.

Sentiment

Constructive, operationally focused +0.12

26 source items · 66% editorial confidence

Must read

FactSet decomposes the Mag 7’s 118.5% aggregate Q2 earnings growth and 66.2% surprise: excluding Alphabet’s $98 billion and Amazon’s $53.4 billion investment gains, growth falls to 43.2% and the surprise to 4.4%. The clean comparison matters because analysts expect the other 493 S&P 500 companies to grow earnings 26.8% in Q4, ahead of the Mag 7’s 23.2%; the headline therefore overstates operating divergence.

Hernández de Cos evaluates stablecoins and tokenised deposits against three monetary properties—singleness, interoperability and integrity—and argues that tokenised deposits fit the two-tier system better. The speech earns its length by tracing the macro-financial mechanism: stablecoin substitution can lift bank funding costs, shift credit to more procyclical non-banks and accelerate digital dollarisation, while tokenised deposits still require common settlement rails, governance, legal clarity and a careful migration path.

Signals

The 2-year Treasury note yield rose 0.118 to 4.348%, which is the second highest level of the year and the largest one-day move since March 12. Futures markets price in a 60% chance of a Sept hike from the Fed and a 90% chance of a hike this year.

The two-year Treasury note yield rose 11.8 basis points to 4.348%, its largest one-day move since March 12, while futures priced a 60% chance of a September hike and 90% chance of one this year. That is a market-pricing consequence of the inflation-first message, not a Federal Reserve commitment or a verified forecast.

updates 2026-08-28 · prior evidence ↗

Commercial and industrial loans at U.S. commercial banks increased $11.8 billion to $2.95 trillion in the week ending August 19. The latest weekly vintage is a useful credit pulse, but one increase does not establish an acceleration or a durable lending trend.

Sentiment

Cautious, policy-sensitive -0.05

126 source items · 61% editorial confidence