All series

Series · 3 parts

Measuring Agent Skills

Three measurements of how AI agents actually use a skill library: why a list failed, why subagents never browsed it, and what Claude Code's native plugin evals do and do not prove.

  1. Part 1We gave our AI agent 500 skills. It needed a graph, not a list.Flat lists of AI agent skills fail silently as they grow. Every failure we hit was a relationship failure, so we built a generated graph, and the case for generating it came from our own code drifting within hours.Jim Deola · August 9, 2026
  2. Part 2Your Subagents Aren't Using Your SkillsWe counted how often our AI subagents invoked skills from a several-hundred-skill library: zero of fifteen dispatches. Here is why cold-started agents never browse the library, the three injection levers that work, the spike methodology that proved the hook surface, and the log-first measurement instrument now watching every dispatch.Jim Deola · August 26, 2026
  3. Part 3Claude Code Plugin Evals: The Traps the Docs Don't MentionWe ran Claude Code's new plugin eval harness on a real plugin: 8 cases, 46 agent runs, about $12. Here are the three ways the results JSON misleads you, what the eval sandbox can and cannot reach, and the five-line hook that fixes the plugin-root variable Bash never sees.Jim Deola · September 16, 2026

See exactly where your business
runs through you.

The next step is a 30-minute call.

Book a 30-minute call and we'll map where the business depends on you — and what it looks like once a system you own runs the routine. You leave with a plan, whether you hire us or not.

No pitch — you leave with a plan. We'll tell you exactly what we'd do, whether you hire us or not.