Back to Resources

Series

Rhize Media resource articles grouped into ordered series: how our plugins, workflows and technical stack change over time, one part at a time.

All series

Pick a story

3 parts · updated September 16, 2026

Measuring Agent Skills

Three measurements of how AI agents actually use a skill library: why a list failed, why subagents never browsed it, and what Claude Code's native plugin evals do and do not prove.

Open the series
In this series · 3 partsShow the parts
  1. Part 1We gave our AI agent 500 skills. It needed a graph, not a list.Flat lists of AI agent skills fail silently as they grow. Every failure we hit was a relationship failure, so we built a generated graph, and the case for generating it came from our own code drifting within hours.Jim Deola · August 9, 2026
  2. Part 2Your Subagents Aren't Using Your SkillsWe counted how often our AI subagents invoked skills from a several-hundred-skill library: zero of fifteen dispatches. Here is why cold-started agents never browse the library, the three injection levers that work, the spike methodology that proved the hook surface, and the log-first measurement instrument now watching every dispatch.Jim Deola · August 26, 2026
  3. Part 3Claude Code Plugin Evals: The Traps the Docs Don't MentionWe ran Claude Code's new plugin eval harness on a real plugin: 8 cases, 46 agent runs, about $12. Here are the three ways the results JSON misleads you, what the eval sandbox can and cannot reach, and the five-line hook that fixes the plugin-root variable Bash never sees.Jim Deola · September 16, 2026

2 parts · updated August 26, 2026

Procedural Memory

How we taught our AI systems to keep the code they already wrote: the registry, the first paired benchmark, and what each measurement changed about the method.

Open the series
In this series · 2 partsShow the parts
  1. Part 1AI Agent Memory: Why Our Systems Stopped Re-Solving the Same ProblemsMost AI automations re-derive the same work every run. Here is how procedural memory, storing the working code rather than a description of it, makes AI agent workflows faster, cheaper, and more predictable.Jim Deola · August 23, 2026
  2. Part 2The Benchmark Pointed at ItselfThe first paired benchmark of our procedural-memory registry measured a deterministic share under 0.5% on one workload and about 55% on another, caught its own row-append step failing silently, and ran under two conditions that rule out any arm-versus-arm verdict. Here is what held up, what did not, and the six variables the next cohort holds fixed.Jim Deola · August 26, 2026

See exactly where your business
runs through you.

The next step is a 30-minute call.

Book a 30-minute call and we'll map where the business depends on you — and what it looks like once a system you own runs the routine. You leave with a plan, whether you hire us or not.

No pitch — you leave with a plan. We'll tell you exactly what we'd do, whether you hire us or not.