Why long-horizon agents fail at hour three
Nothing dramatic happens. The agent keeps producing plausible output, the tests keep passing, and the work quietly stops being about the task.
What we tried, what broke, and what the numbers actually said. Written by the people running the tracks.
Nothing dramatic happens. The agent keeps producing plausible output, the tests keep passing, and the work quietly stops being about the task.
Simulations agree with themselves. The interesting question is where the model and the board disagree, and how fast you find out.
Above a certain latency the loop stops feeling like your own body. Below it, users stop thinking about the interface at all.
The migration deadline is not when quantum computers arrive. It is now minus however long your data stays sensitive.
For walking through a captured space, the mesh pipeline is now the slow path that looks worse.
Nobody has to be malicious. A stuck sensor produces confident wrong data, which is indistinguishable from an adversary.
Most of what we learned was about coordination cost, and none of it was about the model.
The result you cannot rebuild in six months is a rumor. Almost every time we failed to rebuild one, the cause was environment, not science.