The engineering write-up defines a non-autoregressive structured decision model, reports 70–500 ms latency and $0.042 per million input tokens, and discloses benchmark and zero-schema-error limitations.
daily notes
previous notes →Must read
Signals
Biologists use specialized open-source models for tasks like modeling the structure of molecular systems, designing drug-like molecules, and predicting the effects of genetic mutations. But these models are often expensive to run, potentially limiting their impact. In our…
Biologists use specialized open-source models for tasks like modeling the structure of molecular systems, designing drug-like molecules, and predicting the effects of genetic mutations. But these models are often expensive to run, potentially limiting their impact. In our latest Science Blog, we share how Claude was able to optimize inference for more than 30 open-source models, making them 4x faster on average, partly by writing custom software for GPUs. We’re open sourcing all of the optimization code. Read more: https://t.co/qiuN1jpgpA
discussion2 selected replies
Anthropic says Claude optimized inference for more than 30 open-source models, making them four times faster on average, partly through custom GPU software, and says the optimization code is open source. The post does not specify the full workload mix.
We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key…
We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone. https://t.co/yUf6OpJD7c
Z.ai says GLM-5.3 helped take its inference system from first successful run to production readiness in under two weeks, with end-to-end throughput tripling. It attributes the result to local tests, traces, microbenchmarks, and end-to-end measurement.
Sentiment
91 source items · 55% editorial confidence
Must read
The paper documents three measurement divergences across Bitcoin, Ethereum, and Tron and recommends granular, assumption-explicit, disaggregated estimates.
Signals
The Japanese central bank has lifted interest rates by 0.25 percentage points, accelerating its schedule of monetary policy normalisation under mounting pressure from Washington. https://t.co/npotI4vXCg? https://t.co/8howYf3sc0
The post reports that Japan's central bank raised rates by 0.25 percentage points and accelerated monetary-policy normalization under pressure from Washington. It does not state the resulting policy path or market reaction.
Sentiment
466 source items · 55% editorial confidence
To show what these optimizations make possible, we’re partnering with Adaptyv Bio on a protein design competition. Together, we’ll be experimentally validating over 5,000 designs. We're providing up to $1 million in Claude credits plus funding alongside Adaptyv for experimental validation. Modal is contributing up to $250,000 in compute and Twist Bioscience is providing DNA. Learn more on Adaptyv’s Proteinbase: https://t.co/KnjkWbH7zh And sign up for the competition here: https://t.co/ytR2TAC3Vt
You can find all of the code on GitHub: https://t.co/MmPk9hIZpx And the full results in our technical report: https://t.co/HSPEg5Gpsf