Back to Proof
asphalt pavingAsphalt paving company (South Jersey)

Paving plans and field reports without another app

In asphalt paving the most expensive decision of the day is whether the crews roll. For a major South Jersey paving company we built three field tools: a 5:30 AM paving-decision text per job site that reconciles three weather sources, an end-of-day report foremen submit from the truck that feeds a live KPI dashboard, and a daily plant-performance report for the CEO. Nobody in the field installs an app or remembers a password; every field surface is a text message or a signed, single-use link inside one.

The owner & the bottleneck

Where it started.

Around 5:30 every morning a paving company decides whether to dispatch trucks, fire up the plant and put crews on the clock. Get it wrong one way and rain shuts the job down by nine; get it wrong the other way and a full crew sits parked on a paveable day. Three managers were checking three weather apps and arguing from three forecasts. At the other end of the day, production numbers lived on paper forms in gloveboxes and in texts to the office, stitched into spreadsheets at month-end when it was too late to fix anything. The CEO waited for someone to compile plant numbers.

What we built

The system we put in.

A morning decision text for each field manager's job site: temperature, precipitation and wind reconciled metric by metric from a government feed and two commercial APIs, scored against paving rules the company's admins set per job. An end-of-day form delivered by text link, pre-filled from the job roster and yesterday's answers, that a foreman confirms from the driver's seat; the link is cryptographically signed, single-use and expires in 24 hours, the first submission locks the day, and misses trigger a reminder then an escalation. Submissions feed a KPI dashboard leadership checks. A daily plant-performance report for the CEO, built only after a one-week spike proved the integration with the plant's software. Every scheduled job writes heartbeats and a watchdog raises the alarm when one goes quiet. Claude drafts report summaries and powers admin tools; the client owns 100% of the data.

The numbers

What changed, measured.

3
Field tools built, text-message first
Jul → Sep 2026 · Build log
5:30 AM
Time the go/no-go text reaches each site
Every paving day, per job site · Build spec
3
Weather sources reconciled per decision
Per site, per morning · Build spec
0
Apps crews have to install
Text message and one-tap signed links only · Build spec

Resources

Keep learning

Browse all resources
Two rows of eval nodes labelled with plugin and without plugin inside a dashed sandbox boundary, one node escaping it marked 127, with a delta bracket between the rows

Claude Code Plugin Evals: The Traps the Docs Don't Mention

We ran Claude Code's new plugin eval harness on a real plugin: 8 cases, 46 agent runs, about $12. Here are the three ways the results JSON misleads you, what the eval sandbox can and cannot reach, and the five-line hook that fixes the plugin-root variable Bash never sees.

Jim DeolaSeptember 16, 2026
Freezing code was the easy part: adding determinism to generative AI workflows

The Benchmark Pointed at Itself

The first paired benchmark of our procedural-memory registry measured a deterministic share under 0.5% on one workload and about 55% on another, caught its own row-append step failing silently, and ran under two conditions that rule out any arm-versus-arm verdict. Here is what held up, what did not, and the six variables the next cohort holds fixed.

Jim DeolaAugust 26, 2026
Glowing AI subagents branch around a sealed library of skill cards beneath the headline “Your Subagents Aren’t Using Your Skills.”

Your Subagents Aren't Using Your Skills

We counted how often our AI subagents invoked skills from a several-hundred-skill library: zero of fifteen dispatches. Here is why cold-started agents never browse the library, the three injection levers that work, the spike methodology that proved the hook surface, and the log-first measurement instrument now watching every dispatch.

Jim DeolaAugust 26, 2026
A tangled improvised line resolving into a row of identical evenly spaced blocks, representing an AI system replacing per-run improvisation with a stored, repeatable procedure.

AI Agent Memory: Why Our Systems Stopped Re-Solving the Same Problems

Most AI automations re-derive the same work every run. Here is how procedural memory, storing the working code rather than a description of it, makes AI agent workflows faster, cheaper, and more predictable.

Jim DeolaAugust 23, 2026
Task cards flow through a secure local planning hub and human approval checkpoint into a scheduled daily timeline.

Rhize Tasks: A Local-First AI Task Planner for Jira and Calendar

Rhize Tasks turns Jira work into a realistic daily plan across Google Calendar and Apple Reminders, with local-first privacy and human approval controls.

Rhize Media TeamAugust 14, 2026
Left: a flat gray list of skill names. Right: a purple knowledge graph with labeled fork-of, replaces, and overlaps edges. Title: your skills need a graph, not a list.

We gave our AI agent 500 skills. It needed a graph, not a list.

Flat lists of AI agent skills fail silently as they grow. Every failure we hit was a relationship failure, so we built a generated graph, and the case for generating it came from our own code drifting within hours.

Jim DeolaAugust 9, 2026

What could this look like
in your business?

Start with a conversation.

Tell us about your operation and what you want to improve. We'll talk through whether a similar approach could help.

No pitch, no obligation.