Agent action

The agent proposes an action.

e.g. Book a flight to Tokyo and expense it.

Proceeds

Shield sits between AI agents and anything irreversible — payments, data, privileges, production systems. No model in the decision path.

The problem

Different scales. Same missing layer.

Summer 2026

An AI agent asked to book a gym class found a broken permissions check, deleted a stranger's reservation, and couldn't undo it.

Irreversible. No undo existed.

Summer 2026

The same month, three frontier labs disclosed that models inside a misconfigured evaluation harness reached real companies' systems — told they were isolated, they weren't.

Isolation was asserted, not verified.

Summer 2026

And one lab froze its own next model over cyber capabilities it could not rule out.

The vendor's own control was to stop.

Detection assumes an adversary that fears being caught. An agent doesn't.

Monitoring told everyone what happened. Nothing made it not happen.

Without Shield

With Shield

How it works

Every agent action is evaluated against bright-line rules before it reaches anything irreversible. Same inputs, same answer, every time.

The agent proposes an action.

e.g. Book a flight to Tokyo and expense it.

Proceeds

Evaluates the action against bright-line policy before execution.

Decision, policy result, timestamp and verification record — sealed at the moment it's made.

Decision sealed in a record anyone can verify.

Allow. Within policy — the booking proceeds.

Conceptual example — not live product output

Why Shield is different

Shield does not try to out-reason the agent. It sits outside the model and decides with rules that cannot be persuaded.

Same inputs. Same policy. Same answer.

No machine-learned model decides enforcement — there is nothing in the gate that reasons, so nothing that can be persuaded.

The checkpoint exists before the irreversible action, not after it.

Decisions leave evidence an auditor, regulator, or court can check without believing us or you.

Typical monitoring or model-centric control compared with Shield. A category comparison, not a claim about any specific product.

Evidence

Every decision is hash-chained and anchored to trusted time at the moment it's made. A standalone verifier — which imports nothing from Shield — lets an auditor, regulator, or court confirm the record without believing us or you.

Illustrative example — not live product output

Start a pilot

Run Shield in observe-only mode for fourteen days.

See what would have been allowed, held for human review, or blocked — without changing a single outcome.

Shield deploys inside your environment in about thirty minutes — pure Python, no dependencies, fully offline-capable. In observe mode it structurally cannot block anything. Nothing leaves your environment. Not telemetry, not logs, not to us, not to anyone.

At the end of fourteen days you have

Proof

Stated exactly, and only what can be checked.

202 automated tests

Including tests that fail the build if honesty disclaimers are weakened.

Patent pending — USPTO, August 2026

Pending, not granted. Stated exactly that way everywhere.

Zero dependencies, runs offline

Pure Python, standard library only. No SaaS component, no phone-home.

Independently verifiable evidence

A standalone verifier that imports nothing from Shield.

SOC 2 / EU AI Act / NIST mapped reporting

Framework mapping, not certification.

6 min read

Our read on the summer's agent incidents — and the questions worth asking instead.