Back to Resources

What Laya Taught Us About AI Workflow Decisions

Shadow routing, AI-reviewed labels, and why logged events are not tasks

Rhize Media TeamOctober 6, 2026
BuilderDeveloper WorkflowAgent SkillsArticle
On this page

TL;DR

Laya is an open-source local AI decision model that returns bounded answers: a choice, a score, or a yes/no, without calling an external service. In RHIZE Media's pilot, Laya only advised a content/general/none workflow route in shadow mode; the existing agent route (Arm A) was the only thing that executed. The current evidence shows better visibility and measurement of how routing decisions get made, not a proven improvement in accuracy or task quality. Several claims, including whether Laya's advice is better, remain unsupported pending a larger, balanced study.

What Laya Is, Before the Pilot

Independent of any pilot result, Laya is worth understanding on its own terms. It's an open-source local typed-decision model: you give it a question, and it gives back a bounded answer type, a discrete choice from a known set, a numeric score, or a yes/no. This typed-decision pattern, where outputs are constrained to known shapes rather than free text, is documented in RHIZE's typed decision layer reference. Because it runs locally, it doesn't require a network round-trip to an external API, which means lower latency and no outbound call carrying your data. Bounded, typed outputs also make decisions easier to audit after the fact, since "the model picked B with score 0.7" is a concrete, loggable fact rather than a paragraph of free text to parse. Those are useful properties for any system that wants a fast, inspectable gatekeeper step, before you even get to the question of whether a specific pilot proved anything.

Shadow Advice vs. Executed Decisions: What Arm B Did

Think of it like a co-pilot calling out a turn while someone else keeps both hands on the wheel. The suggestion is logged. It is not followed.

In the pilot, Arm A is the existing agent route, the one that runs tasks today. Arm B is Laya's recommendation of a content/general/none workflow, generated in shadow mode only. Per the decision-pilot documentation, Laya's advice in this pilot never authorized publication, triggered a tool call, bypassed a safety check, or caused a deployment. It was recorded alongside Arm A's actual route so the two could later be compared, but comparison is not the same as control. Nothing in this pilot let Laya's opinion change what happened.

Reading the Numbers: Raw Events Are Not Independent Tasks

A raw "event" can be one logged step inside a task. The 1,962 v2 events below are not 1,962 independent decisions; the comparable set is much smaller.

The October 6 v2 snapshot counts 1,962 raw events across current, historical, and missing-context sources. The current-source digest (e8c40f12...75f9bd) has 79 events; historical source changes and missing-context events remain visible but are held out of current comparisons. There are 43 eligible task roots, all scored: 39 on Codex and four on Claude. Only 32 have comparable Arm A and Arm B routes, all on Codex; 30 of those differ. That is a useful review queue, not an accuracy rate. We still have zero human correctness labels, zero accepted tasks, and no complete coding-agent usage measurement.

Table 1: Pilot Scale at a Glance

MetricCount
Raw v2 events (all sources)1,962
Current-source events79
Eligible task roots43
Scored task roots43
Codex-eligible39
Claude-eligible4
Comparable A/B routes (Codex only)32
Differing routes30
Accepted tasks0
Human correctness labels0

Historical and missing-context events remain visible in the report and are intentionally not pooled with current decisions.

AI-Reviewed Labels Are Not Human Ground Truth

An AI-reviewed label means a model judged the outcome, not a person. It's a useful, fast signal for development work. It is not a validated ground-truth benchmark.

The pilot separates human_adjudicated from ai_model_reviewed labels. As of October 6, all 38 labels are AI-reviewed; 33 have an explicitly judged route, and 16 belong to the current source. Across all 38, content has two labels, general 28, and none eight. This uneven coverage describes what got labeled, not which route was correct. The latest scheduled labeler run (20261006T113006Z-2a216b44) selected 17 cases: four were labeled, nine remained unresolved, and four lacked sufficient context. Claude Sonnet 5.5 and Codex 5.6 Sol annotated; Claude Fable 5.1 reviewed. These labels help development research but are not human ground truth or a validated performance benchmark.

Taxonomy: Family, Phase, Risk, and Why a Family Doesn't Pick the Route

Knowing a task's "family" tells you what kind of work it is, not which workflow route handled it. Those are recorded as separate facts, on purpose.

The pilot introduces a finer-grained workflow taxonomy with nine families (feature delivery, defect resolution, code health, platform operations, content/growth, research/analysis, knowledge management, coordination, and direct response), plus an immediate phase, the affected areas, and a risk stratum. All of that is recorded separately from the content/general/none route choice. That separation matters: a "feature delivery" task doesn't automatically route to "general," and a "defect resolution" task doesn't automatically route to "none." The taxonomy gives reviewers a richer description of what a task is; it doesn't dictate or explain why Arm A or Arm B picked the route it did.

Check Passed, Task Accepted: Two Different Things

A command finishing without error is like a single part passing its own certification. It tells you that one piece worked, not that the whole assembly was accepted into the final build.

An automatic or command-level check succeeding is evidence that a command ran and returned cleanly. It is not evidence that a human looked at the resulting task and accepted it. That distinction is why, even with 1,962 raw v2 events on file, accepted tasks still sit at zero. Nobody has yet formally signed off that a task's output was good, correct, or usable. Conflating "the check passed" with "the task is done" would overstate what the pilot has shown.

What We Can Claim Today

The supportable, positive finding from this round is about measurement quality, not outcomes: the pilot now captures tasks with more context, separates consultation evidence from route-match evidence from execution evidence, produces tighter review packets for human reviewers, applies a finer taxonomy, and records explicit abstentions instead of silently dropping unclear cases. That is real progress in visibility.

What it does not support: a causal improvement in task quality, cost, or speed, and no claim that Laya's routing is more accurate than the existing agent route. With 33 research-usable explicit labels against a floor of 200, plus sparse routing, family, and risk coverage, claims about Laya's performance remain unsupported.

The Next Experiment, in Plain English

The honest next step is a proper, controlled comparison, not a bigger pile of shadow logs. That means balanced, source-matched labels instead of the current lopsided mix; a closed holdout set that isn't touched during analysis; and a run where Arm B is selected and followed on matched real tasks, under controlled conditions, rather than only logged in the background. That run should record whether the resulting task was accepted, how much rework it needed, how long it took, and whether it involved complete coding-agent usage. That is the concrete path from "we can see more clearly" to "we have evidence of improvement."

Table 2: What Would Change Between Now and a Research-Grade Study

ElementCurrent PilotPlanned Study
Arm B roleLogged advice only (shadow)Selected and followed on matched tasks
Label sourceAI-reviewed onlyBalanced, source-matched labels
Holdout setNoneClosed holdout, untouched during analysis
Outcome trackedRoute match/differenceAcceptance, rework, time, full agent usage
Label count38 total; 33 explicitAt least 200 eligible, source-bound explicit labels plus coverage gates

Sources and Further Reading

These figures use RHIZE's October 6 aggregate pilot snapshot. The Content Engine comparison used a frozen October 4 packet; this publication update changes the figures without counting it as a new graph run. Task transcripts and credentials are excluded. Public implementation and evaluation documents are linked below.

Watch for RHIZE Media's next write-up once the balanced, holdout-controlled study runs Arm B on real tasks.

Want this working in your business — not just on paper?

Book a 30-min call to talk through your workflow, ask your questions, and discuss a useful next step. No pitch, no obligation.

Book a 30-min call

Get insights like this in your inbox

Join our newsletter for actionable SEO, marketing, and growth strategies — no fluff, just results.

No spam. Unsubscribe anytime.

See exactly where your business
runs through you.

The next step is a 30-minute call.

Book a 30-minute call and we'll map where the business depends on you — and what it looks like once a system you own runs the routine. You leave with a plan, whether you hire us or not.

No pitch — you leave with a plan. We'll tell you exactly what we'd do, whether you hire us or not.