
The Benchmark Pointed at Itself
The first paired measurement of a procedural-memory registry against LLM recomposition found the deterministic share was under 0.5% of run time, and that the benchmark's own row-append step had been silently failing. Here is what that taught us about verification, and where determinism actually belongs in an agent workflow.




























