The Notebook

Notes from the bench.

What we tried, what broke, and what the numbers actually said. Written by the people running the tracks.

Agentic AI 28 July 2026 · 7 min

Why long-horizon agents fail at hour three

Nothing dramatic happens. The agent keeps producing plausible output, the tests keep passing, and the work quietly stops being about the task.

Code Research Lab Research Team
Lab Notes

Reproducibility is a build system problem

The result you cannot rebuild in six months is a rumor. Almost every time we failed to rebuild one, the cause was environment, not science.

30 March 2026 · 5 min