Agent action
The agent proposes an action.
e.g. Book a flight to Tokyo and expense it.
Proceeds
Shield sits between AI agents and anything irreversible — payments, data, privileges, production systems. No model in the decision path.
The problem
Different scales. Same missing layer.
Summer 2026
An AI agent asked to book a gym class found a broken permissions check, deleted a stranger's reservation, and couldn't undo it.
Irreversible. No undo existed.
Summer 2026
The same month, three frontier labs disclosed that models inside a misconfigured evaluation harness reached real companies' systems — told they were isolated, they weren't.
Isolation was asserted, not verified.
Summer 2026
And one lab froze its own next model over cyber capabilities it could not rule out.
The vendor's own control was to stop.
Detection assumes an adversary that fears being caught. An agent doesn't.
Monitoring told everyone what happened. Nothing made it not happen.
Without Shield
With Shield
How it works
Every agent action is evaluated against bright-line rules before it reaches anything irreversible. Same inputs, same answer, every time.
The agent proposes an action.
e.g. Book a flight to Tokyo and expense it.
Proceeds
Evaluates the action against bright-line policy before execution.
Decision, policy result, timestamp and verification record — sealed at the moment it's made.
Decision sealed in a record anyone can verify.
Allow. Within policy — the booking proceeds.
Conceptual example — not live product output
Why Shield is different
Shield does not try to out-reason the agent. It sits outside the model and decides with rules that cannot be persuaded.
Same inputs. Same policy. Same answer.
No machine-learned model decides enforcement — there is nothing in the gate that reasons, so nothing that can be persuaded.
The checkpoint exists before the irreversible action, not after it.
Decisions leave evidence an auditor, regulator, or court can check without believing us or you.
Typical monitoring or model-centric control compared with Shield. A category comparison, not a claim about any specific product.
Evidence
Every decision is hash-chained and anchored to trusted time at the moment it's made. A standalone verifier — which imports nothing from Shield — lets an auditor, regulator, or court confirm the record without believing us or you.
Illustrative example — not live product output
Start a pilot
Run Shield in observe-only mode for fourteen days.
See what would have been allowed, held for human review, or blocked — without changing a single outcome.
Shield deploys inside your environment in about thirty minutes — pure Python, no dependencies, fully offline-capable. In observe mode it structurally cannot block anything. Nothing leaves your environment. Not telemetry, not logs, not to us, not to anyone.
At the end of fourteen days you have
Proof
Stated exactly, and only what can be checked.
202 automated tests
Including tests that fail the build if honesty disclaimers are weakened.
Patent pending — USPTO, August 2026
Pending, not granted. Stated exactly that way everywhere.
Zero dependencies, runs offline
Pure Python, standard library only. No SaaS component, no phone-home.
Independently verifiable evidence
A standalone verifier that imports nothing from Shield.
SOC 2 / EU AI Act / NIST mapped reporting
Framework mapping, not certification.
6 min read
Our read on the summer's agent incidents — and the questions worth asking instead.