Mollick uses the Hugging Face Incident to argue for a Twilight Factory: agents can execute large amounts of work, but they should escalate decision points where human context, taste or risk judgment changes the direction. The essay's durable value is the operating model, not a new incident claim; it shifts oversight from reviewing every action to designing explicit boundaries where agents must ask.
daily notes
previous notes →Must read
Apodex has launched Apodex 1.1, a proprietary model scoring 44 on the Artificial Analysis Intelligence Index with strong performance in agentic tasks compared to models in its intelligence tier Apodex 1.1 is Apodex's first model on Artificial Analysis. The lab has previously…
Apodex has launched Apodex 1.1, a proprietary model scoring 44 on the Artificial Analysis Intelligence Index with strong performance in agentic tasks compared to models in its intelligence tier Apodex 1.1 is Apodex's first model on Artificial Analysis. The lab has previously launched Apodex 1.0 and Apodex 1.0 mini. At 44 on the Intelligence Index, Apodex 1.1 sits in a similar tier alongside Kimi K2.6 (45), and MiniMax-M3 (45), and Inkling (42). Within that group, Apodex 1.1 stands out on agentic and knowledge work evaluations, but demonstrates trade-offs on knowledge reliability and frontier academic reasoning. Key results: ➤ Apodex 1.1 demonstrates a key strength in agentic workflows. On GDPval-AA v2, our real-world agentic work benchmark, Apodex 1.1 achieves an Elo of 1348, which places it ahead of DeepSeek V4 Pro (1333), Qwen3.7 Max (1308), and Kimi K2.6 (1202). On TerminalBench v2.1, our agentic coding and terminal use benchmark, it also performs strongly with a strong 70% score, ahead of Kimi K2.6 (66%) and just behind Qwen3.7 Max (75%). ➤ Apodex 1.1 uses ~17k output tokens per task on average across the Artificial Analysis Intelligence Index. This makes the model more token efficient than DeepSeek V4 Pro (16,842 tokens), but less token efficient than other models with a similar Intelligence Index score, such as Qwen3.7 Max (9,391 tokens) and MiniMax-M3 (8,133 tokens). ➤ Apodex 1.1 costs ~$0.05 per Intelligence Index Task, making it relatively attractive compared to peer models in its intelligence tier. This is primarily driven by a cheap pricing but is slightly offset by the higher tokens per task. At ~$0.05, Apodex 1.1 is cheaper than Qwen3.7 Max (~$0.07) and Kimi K2.6 (~$0.06). Even though Apodex 1.1 is not on the Cost vs. Intelligence Pareto frontier, it lands in the most attractive quadrant. ➤ Apodex 1.1 demonstrates modest performance on knowledge accuracy and reliability, scoring -21.9 on AA-Omniscience. The model achieves 32% accuracy on individual questions, with a 78.4% hallucination rate. It attempts to answer 87% of questions rather than declining to respond. Additional model details: ➤ Context window: 256K tokens ➤ Pricing: $0.30 / $3.00 per 1M input/output tokens, with a $0.03 cache-hit price ➤ Availability: Apodex first party API
Artificial Analysis measured Apodex 1.1 at 1,348 Elo on GDPval-AA v2 and 70% on TerminalBench v2.1 while estimating about $0.05 per Intelligence Index task. The same evaluation found only 32% question accuracy and a 78.4% hallucination rate on AA-Omniscience. That combination makes it a plausible low-cost agent model only when the workflow can verify outputs; it is a poor default for knowledge-sensitive work.
Tailcat packages WireGuard encryption, NAT traversal and DERP fallback as an open-source Go library and CLI without accounts, admins, visible IPs or a Tailscale control plane. The article explains the key exchange, DERP rendezvous, direct-path fallback and userspace TCP design in enough detail to judge where it fits. Its sharpest use is giving short-lived or untrusted agent sandboxes encrypted connectivity without enrolling them in a durable network identity system.
Signals
For clarity, while both are called 20X, in Codex they apply specifically to weekly usage limits. And we also don't have 5h limits for both Pro plans. The Pro 20X is quite precisely 20X the usage of the Plus subscription, so it does exactly what it says on the tin.
The clarification narrows the plan comparison: Codex's 20X multiplier is specifically a weekly allowance, and neither Pro plan has a five-hour limit. Capacity planning should therefore compare weekly work volume rather than assume every usage window scales by the same factor.
Google’s 330M-parameter model is pretrained on more than 1T time points and jointly handles multiple targets plus historical and known-future covariates. Its non-autoregressive decode produces the full horizon and nine quantiles in one forward pass, making it a practical zero-shot baseline when iterative decoding latency or compounding forecast error is the constraint.
Sentiment
54 source items · 68% editorial confidence
Must read
The authors separate the level of expected inflation from disagreement, individual uncertainty, skewness and tail risk, then show how survey, market and model measures observe different populations and information sets. Their practical proposal is to track probability mass across deflation, near-target and high-inflation regions with standard macro tools. That turns an average expectation into a map of where policy credibility is actually breaking.
The authors find no average linear sentiment response to unexpected policy-rate changes, then recover a state-dependent split: a one-standard-deviation tightening surprise is associated with a 1.43% sentiment increase after low inflation but a 1.38% decline after high inflation. The 2004–2024 design, robustness checks and information-effect interpretation make the full paper useful for deciding when a rate move communicates strength and when it mainly raises borrowing pain.
Signals
Draft IPO documents reviewed by the Wall Street Journal put the estimated value of warrants issued to OpenAI at $5.5 billion as SB Energy sought the company as a data-center tenant. The number is not final until a filing confirms it, but it makes the tenant incentive explicit and should be separated from the data-center unit's operating economics.
The Financial Times puts the quarterly accounting boost from Big Tech stakes in other AI companies above $160 billion. That quantifies the sector-wide scale behind the earlier Mag 7 warning: reported profit growth can rise with private-company marks even when no operating cash enters the business.
updates 2026-08-29 · prior evidence ↗
BYD's first-half overseas revenue rose 34% to 53% of total revenue while Greater China revenue fell 31%, according to Bloomberg's account of the results. The crossover turns exports from a growth option into the company's main revenue geography, increasing exposure to tariffs, local manufacturing rules and overseas distribution execution.
The Dallas Fed production index rose six points to 16.1 in its Aug. 18–26 survey, showing faster factory output growth. The broader Texas business survey still reports net margin compression and weaker pricing power, so the production acceleration is evidence of firmer activity rather than a clean profit rebound.
Together AI’s Humain partnership covers 250MW and 120,000 accelerators, with the startup offering revenue share in exchange for capacity and projecting $5B in annual data-center revenue. The structure shows power access becoming financing consideration; the revenue figure is management’s expectation, not realized sales.
Sentiment
370 source items · 73% editorial confidence