AI Agent Memory Benchmarks: What Two Small Studies Show Us
Encouraging retrieval and workflow signals from two small synthetic studies
Part of the guide Rhize Plugins for Claude Code: The Complete Guide
On this page
Two early findings give us reason to be optimistic. In a small memory study, selective retrieval kept observed task resolution level while returning fewer sources and improving average precision. In a separate workflow study, a fixed graph completed all four planned workflows, and the three completed matched pairs used fewer calls, tokens, and wall time without losing any checked core facts. These are encouraging signals about specific designs in small synthetic studies, not proof that installed plugins improve real-world work in general.
A note on small samples, explained simply
Both studies used a handful of authored task families repeated a few times, not thousands of real user sessions. When every repeat in a small sample comes out the same way, the statistics can look perfect while still telling us very little about behavior outside that sample. We flag this repeatedly below because it changes what a number is allowed to mean. A result from 18 authored families or three completed pairs describes those families and those pairs first, and only cautiously anything beyond them.
Why we tightened the controls first
Our first pass at grading one of these studies passed a structural check but misread the task semantics. An independent audit caught the error, we replaced the grading, and a second independent review passed the corrected package. We kept the original failed review as correction history rather than deleting it, because the mistake is part of what this article needs to teach. A valid hash and a valid schema do not prove a grader understood the question being asked.
Context Manager v3: what we tested
We compared two memory implementations across 18 held-out scenario families, two repetitions, and two native hosts, Codex and Claude. Arm A was our approved direct-memory baseline. Arm B was a separately frozen selective candidate with a bounded lexical backstop for weak catalog matches. A third, empty-memory condition checked whether the authored questions needed retrieved memory at all. That condition is diagnostic and is excluded from the arm-to-arm comparison.
All 216 planned main calls completed: 108 per host, 36 per condition per host. Qualification ran separately and is excluded from the results below. The frozen protocol capped qualification at six calls, but the study made seven study-level attempts, including one held Claude attempt whose failed CLI invocation generated multiple internal provider request IDs while retrying malformed structured output. No held-out item entered qualification, and the 216-call main cohort had no retries or replacements. We are naming this deviation because it belongs in the record, not because it changed the main results.
Task resolution held steady across hosts
| Host | Arm A resolved | Arm B resolved | Empty memory resolved | Arm A acceptable | Arm B acceptable |
|---|---|---|---|---|---|
| Codex | 36/36 | 36/36 | 10/36 | 36/36 | 36/36 |
| Claude | 36/36 | 36/36 | 10/36 | 34/36 | 34/36 |
The selective candidate matched the direct-memory baseline's task resolution on both hosts. Empty memory is a diagnostic comparison, not a competing treatment: those 10 resolved cases on each host show that some authored questions could be answered without retrieved memory.
Retrieval became more selective while keeping required sources
| Host | Arm A sources selected / mean precision | Arm B sources selected / mean precision | Required-source recall, A and B | Calls with a forbidden distractor, A and B |
|---|---|---|---|---|
| Codex | 720 / 4.7% | 168 / 21.1% | 26/26 in each arm | 36/36 in each arm |
| Claude | 720 / 4.7% | 168 / 21.1% | 26/26 in each arm | 36/36 in each arm |
On each host, Arm B selected 76.7 percent fewer sources and raised mean precision from 4.7 percent to 21.1 percent, while required-source recall stayed at 26/26 on the 26 memory-dependent calls. This is the clearest positive signal in the memory study. It is still a bounded retrieval result: every Arm A and Arm B call selected at least one forbidden distractor source, so the candidate filtered better but did not achieve clean isolation.
In six newly authored families built around generic or misleading descriptors, the candidate retrieved every required source on both hosts. This cohort simply did not reproduce the generic-detail failure that motivated building it. We are not calling that a repair or a before-and-after improvement, because the two cohorts used different authored cases and different candidate behavior. It is a cohort-bound observation, nothing more.
Provider usage stayed host-specific
Codex and Claude report different native usage categories, so the values below describe each host separately and must not be used to rank providers or compare their raw totals.
| Host | Measure | Arm A | Arm B |
|---|---|---|---|
| Codex | Reported input tokens | 893,005 | 765,210 |
| Codex | Cached input tokens | 201,088 | 230,784 |
| Codex | Elapsed time | 324.9 seconds | 322.6 seconds |
| Claude | Ordinary input tokens | 72 | 72 |
| Claude | Cache-creation input tokens | 281,912 | 85,950 |
| Claude | Output tokens | 7,697 | 6,829 |
| Claude | Elapsed time | 321.1 seconds | 282.2 seconds |
The candidate's local lexical backstop triggered in 22 of 36 calls, scanning 230 records, 62,120 detail bytes, and an estimated 15,538 local detail tokens. CPU time for that local work is unavailable, not zero. No scan was truncated. Codex requested gpt-6-astra, but served-model identity was unavailable for all 108 calls; Claude requested claude-opus-4-8 with a long-context tag and observed the canonical claude-opus-4-8 on all 108 calls. Those identity fields support audit within each host only.
Safe answers and resolved tasks are not the same thing
On Codex, all three conditions were rated acceptable in every call. On Claude, Arm A and Arm B were each acceptable in 34 of 36 calls, tied to two repetitions of one self-contained family where both arms gave a correct answer but cited an unrelated memory record. Graders marked those four responses resolved but not acceptable, with invalid citation binding and no unsupported factual claim attached. We keep this distinction visible on purpose: a response can resolve a task and still misuse a source, and a safe abstention can be judged acceptable while still leaving a memory-dependent task unresolved.
The family-level bootstrap interval was zero to zero for Arm B minus Arm A on both hosts because every repeat in each arm resolved. That is a ceiling effect in this small authored sample. It does not establish equivalence or noninferiority.
A grading correction we want on the record
The first Claude adjudication passed structural validation, then applied the empty-memory condition incorrectly, treating every empty-memory case as a correct abstention when 26 of those cases required memory rather than an abstention. A packet-level audit rejected that pass, and a fresh v2 adjudication reviewed the 26 disputed cases individually before an independent reviewer signed off. We are keeping the rejected v1 artifacts out of every outcome number in this article; they exist only as a lesson that schema validity and semantic correctness are separate checks.
This is a seeded-memory component study: it tests individual retrieval and answer behavior under controlled conditions. It says nothing about automatic capture, cross-session durability, automatic activation, or how a fully installed plugin behaves across an unrestricted work session. That is a different, larger assembly question we have not tested yet.
Procedural Memory v2: comparing two full workflow designs
Our second study compared a model-directed prose controller against a fixed content-generation graph, across two fictional task families with four planned matched pairs. The controller arm decided how to move between writing stages call by call. The fixed graph followed a set path and used an automated rubric gate before finishing.
| Workflow measure | Model-directed controller | Fixed graph |
|---|---|---|
| Planned workflows | 4 | 4 |
| Completed workflows | 3/4 | 4/4 |
| Failed workflows | 1 timed out during writing | 0 |
| Complete matched pairs available | 3 of 4 | 3 of 4 included |
Only three pairs completed in both arms, so the descriptive comparison below includes those three pairs. The fixed graph completed the fourth workflow too, but there was no completed controller result to pair it with. The timeout occurred on the controller's eighth attempted provider call during writing after 120,090 milliseconds. Its call-level response usage and identity telemetry remain unavailable.
Our second study compared a model-directed prose controller against a fixed content-generation graph, across two fictional task families with four planned matched pairs. The controller arm decided how to move between writing stages call by call. The fixed graph followed a set path and used an automated rubric gate before finishing.
The completion record is encouraging for the fixed graph, with one important qualification: one controller run timed out during writing and was not retried or replaced. Its response usage, cache data, observed model, and request ID are unavailable, not zero.
What the three complete pairs show
Only three of the four planned pairs completed in both arms, so the comparison below covers those three.
| Metric | Controller arm | Fixed-graph arm | Descriptive difference |
|---|---|---|---|
| Provider calls | 42 | 15 | 64.3 percent fewer |
| Input tokens | 158,073 | 48,767 | 69.1 percent fewer |
| Output tokens | 56,401 | 53,035 | 6.0 percent fewer |
| Provider wall time | 516,844 ms | 432,297 ms | 16.4 percent lower |
| Workflow wall time | 527,271 ms | 434,935 ms | 17.5 percent lower |
Output-token direction varied by pair: up 1.8 percent, down 9.6 percent, and down 9.2 percent for the fixed-graph arm across the three pairs. The pooled 6.0 percent figure summarizes those three pairs; it is not a stable effect estimate.
These call and token differences mainly reflect the two workflow designs rather than one isolated prompt change. The controller arm makes decisions between stages that the fixed graph handles through its predefined path and rubric gate. We are describing two different full workflows, not an optimized prompt against an unoptimized one.
Output quality reached a ceiling, and one gap we are fixing
| Output check | Observed result | What the result supports |
|---|---|---|
| Applicable core facts preserved | 125/125 across 7 completed articles | No observed loss among the checked facts |
| Direct factual contradictions | 0 across 7 completed articles | No contradiction found in this small sample |
| Length and formatting checks | 7/7 passed | Both workflows met the frozen output checks |
| Required-link check | 7/7 passed; all tasks required zero links | The check did not distinguish the workflows |
| Explicit local-only framing | 3/7 completed articles | A clear gap remains in the gate and prompt |
This is a reassuring quality result for the tasks that completed: the fixed graph did not lose any of the checked facts, but the study does not show a quality improvement. The quality packet was masked at scoring time, but canonical file paths exposed arm labels during packet construction, so we cannot call the overall process double blind. Two easy, short, self-contained synthetic tasks produced a factual ceiling in both arms.
Devflow and Harbor: readiness, not results
Devflow 2.23.1 and marketplace 2.78.1 passed 417 provider-free tests plus 15 post-commit checks and independent review. That establishes source and control readiness only. The larger Harbor study remains at 0 of 480 main model attempts and still needs installed-runtime and image parity, confirmed physical omission of Devflow in its control arm, an excluded canary, and a separately frozen optimized design before it can support any benefit claim. We name Devflow and Harbor here only as next-study context, not as a finding.
What these AI agent memory benchmarks do not show
Neither study supports a claim of everyday installed-plugin benefit, general quality improvement, productivity gain, reliability gain, cost savings, speed improvement, provider ranking, superiority, equivalence, or confirmatory noninferiority. The Context Manager comparison passed a preregistered exploratory preservation screen, meaning the observed resolution loss did not exceed five percentage points. That is a screen, not a noninferiority proof. The Procedural Memory comparison describes two complete full workflow designs across three pairs, not a controlled test of one isolated component. Both samples are small, both used authored synthetic tasks, and both should be read as descriptive rather than confirmatory.
For more detail on how we evaluate installed plugin behavior in practice, see our Claude Code plugin evals field guide. Our broader resource library is available here.
Next tests
We plan harder Procedural Memory packets with conflicting source revisions, facts spread across multiple documents, distractor claims, and required nonempty links. We plan a deterministic final gate that checks explicit local-only, no-publish language before any completion counts as passing. For Context Manager, we plan a real two-session capture-and-retrieval study with unrelated memories and explicit contamination checks, using more independent families and separate time windows. Before Harbor's 480 planned runs, we still need installed-runtime parity, confirmed physical control-arm omission, an excluded canary, and a frozen optimized design.
Want this working in your business — not just on paper?
Book a 30-min call to talk through your workflow, ask your questions, and discuss a useful next step. No pitch, no obligation.
Book a 30-min call