In our DeepSWE analysis, Kimi K3 Max came close to Fable 5 xhigh on Pass@1 while costing about one-third as much per rollout.
That resulted in 2.8× more solved tasks per dollar.
For teams running models at scale, the cost per successful task is often the more useful…
In our DeepSWE analysis, Kimi K3 Max came close to Fable 5 xhigh on Pass@1 while costing about one-third as much per rollout.
That resulted in 2.8× more solved tasks per dollar.
For teams running models at scale, the cost per successful task is often the more useful comparison. https://t.co/xqL6g6ZscS
Together's DeepSWE analysis says Kimi K3 Max came close to Fable 5 xhigh on Pass@1 at roughly one-third the rollout cost, yielding 2.8× more solved tasks per dollar. The post does not disclose the full test setup, so treat it as a first-party efficiency claim, not an independent benchmark.
We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for…
We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story. It's kind of janky but fun. But it's a bit mindboggling that the LLM has to place and orchestrate various polygon assets in (x,y,z) coordinates and write code that animates it all, and that it even does anything at all.
I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". There might be a lot more. But I'm excited about creating hyper custom worlds that you can imagine dropping players into, e.g. here to participate in the LoTR story as a spectator NPC, or one of the characters, or etc. Something like an ephemeral GTA of X on demand.
Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them. Here, Opus 5 had to very slowly and painstakingly take screenshots at different points, and it messed up a few times and created a bunch of jank. An example of raw capability (multimodal, gameplay) that I think is still quite lacking.
@cunkpyber Eleven Labs for the audio. LLMs can easily use the APIs (here I did that part manually because I felt picky about the voice).
Opus 5 spent about two hours and a 1M-token budget writing 5,500 lines of procedural Three.js, but it had to inspect slow screenshots and still left errors. Long-horizon stamina improved faster than native visual and game feedback.
Sentiment
stamina rising, self-audit lagging+0.18
41 posts · 74% confidence
Themes
No themes have been published for this stream yet.
Nikkei's 715-company survey puts planned investment at ¥35.6734tn, up 14.2% from prior-year actuals, with AI spending spreading across data centers, semiconductors, and power grids.
Japan said it bought yen on July 31 with the U.S. Treasury to counter excessive volatility and disorderly moves, and it would not hesitate to intervene jointly again. The statement confirms execution after July 31 evidence showed only Treasury preparation.
Taiwan's Q2 GDP grew 12.92% year over year, but Nikkei reports the gains remain concentrated in high-tech clusters; new-home prices near TSMC rose 20% while benefits elsewhere lagged.
@cunkpyber Eleven Labs for the audio. LLMs can easily use the APIs (here I did that part manually because I felt picky about the voice).