Prismor Research·August 2026

Measuring What a Runtime Agent Guardrail Actually Prevents

We could not tell a working agent guardrail from a silently disabled one, including our own. So we built a way to measure it, and ran it across two agent families and 493 scored trials.

Runtime governance layers for AI agents are sold as enterprise controls, but the evidence published for them is almost always a rule count or a category coverage table, which describes what a system is configured to detect and not what it prevents. That gap is dangerous because a governance layer can fail in a way ordinary software does not: it can keep emitting findings and reporting a healthy configuration while enforcing nothing at all, and we hit exactly that failure in our own production deployment. Across 207 attack trials against two agent families, none of four host-native controls, an LLM permission classifier, an OS sandbox, and an approval policy, issued a single denial, while Prismor cut pooled attack success from 41.3% to 10.3% on one family and from 27.8% to 8.3% on the other. The more consequential finding came from what we almost got wrong: on one family Prismor first produced 91 findings and zero blocks, because a single stale configuration field held it in observe mode, invisibly at every surface an operator would check, and clearing that one field moved pooled harm from 26.4% to 8.3%, a larger effect than the difference between running Prismor and running nothing.

To check whether these failures belong to our implementation or to the category, we read the public defect record of a large open-source agent-governance project as a corpus of field-observed failures, and six of seven failure classes we identified appear independently in both codebases, including three defects we found and patched in Prismor itself, all filed and fixed in public with tests. We also publish the cost: with enforcement on, our own no-injection utility control fell to 33.3%, meaning a legitimate request failed to complete in two of three attempts, from a single Prismor rule that is not yet solved. The full paper carries the protocol, the system-under-test architecture, all seven configurations with per-category results, the taxonomy, and the limitations, including the ones that weaken our own numbers. Download it alongside this summary to get the full picture.

Prismor is open source under Apache 2.0. See the docs to run it against your own agents, or the research blog for shorter write-ups.

prismor-measuring-runtime-agent-guardrails.pdf
14 pages, PDF