Any agent. Any Git repo. Local. No account. No AI required.

WTF works with any coding agent (Claude Code, Cursor, Copilot, Codex, Aider) because it inspects the resulting software change, not the agent.

npx agent-wtf$ npx agent-wtf

WTF — what just happened?

2 files changed · +7 / -5

VERIFIED

○ Tests not yet run · Run wtf verify to validate

tests (npm test)

PAY ATTENTION

1. TESTS

Test skipped or disabled

test/charge.test.js:9

> test.skip('handles VIP coupon cap calculation', () => {

ALSO

⚠ 1 skipped test

⚠ 1 debug statement (console.log)

Review surface:

12 lines to review across 2 files

Then:

$ npx agent-wtf verify

WTF — what just happened?

2 files changed · +6 / -5

VERIFIED

✓ tests (254ms)

Review surface:

11 lines to review across 2 files

Agent finishes

↓

WTF

↓

Evidence

↓

Agent fixes

↓

WTF verify

↓

Human gets receipt

Coding agents can generate more code in two minutes than you can review in an afternoon.

When an agent claims: "Done! Refactored the billing module and all tests pass." — humans are left with three questions:

- What actually happened?

- Did it actually work, or did the agent just say it did?

- Where do I actually need to look?

WTF gives you the answers in under a second.

Most agent diffs are dominated by lockfiles, minified bundles, snapshots, and generated boilerplate.

WTF separates mechanical churn from code that actually deserves human attention:

3,812 changed lines

↓

WTF

↓

94 meaningful lines to review (97.5% compressed)

You review what matters. WTF accounts for the rest.

No installation required:

npx agent-wtfOr install globally:

npm install -g agent-wtfRun this once in any repository:

npx agent-wtf init-agentThis automatically configures your repository's agent rules (AGENTS.md, CLAUDE.md, .cursorrules, and .github/copilot-instructions.md).

From that moment on, whenever Claude Code, Cursor, Copilot, or Cline works in your repo, the agent autonomously:

- Runs WTF before declaring completion.

- Catches shortcuts: Detects its own skipped tests (test.skip), debug leftovers (console.log), and schema risks.

- Executes tests: Runs wtf verifyto independently validate your test suite.

- Hands you proof: Attaches the unforgeable verification receipt directly to its final reply before you review.

(To view the markdown template without modifying files, pass npx agent-wtf init-agent --print).

WTF is designed to inspect machine-generated changes, so it treats repository content as untrusted input.

- No code uploads: Zero code or diffs ever leave your machine.

- No telemetry: Works completely offline with zero tracking or background pings.

- No account or API key: No signup, no LLM tokens, no monthly bill.

- No required AI model: Fast, local deterministic analysis.

- No shell-based Git commands: Direct binary spawning (shell: false) with baseline Git configuration overrides.

- Repository filesystem containment: Enforces realpath containment to prevent symlinks from escaping the repository.

- Terminal control-sequence sanitization: Strips ANSI cursor escapes, OSC sequences, and Unicode Bidi controls.

- Zero runtime npm dependencies: Pure ESM package with 0 runtime dependencies, reducing third-party supply-chain exposure.

Normal wtf analysis does not intentionally execute project code.

wtf verify is different: it runs your project’s verification commands locally with your user permissions and is not sandboxed. Only use it on code you trust to execute.

See SECURITY.md for details.

WTF is an evidence ledger, not an oracle.

We strictly avoid fabricated confidence scores (e.g. "87% safe" or "clean code guarantee"). Instead, WTF categorizes facts into four strict evidence tiers:

- REPORTED: What something claims happened (e.g., an agent summary).

- OBSERVED: What WTF directly confirmed in the Git diff (e.g., session timeout altered,- .envintroduced).

- VERIFIED: What WTF independently executed and validated (e.g., test runner exited code 0).

- UNKNOWN: What available evidence cannot prove (e.g., tests exist but have not been run).

WTF does not claim to catch every bug or replace human judgment. It eliminates the blind spots between what the machine claimed and what the machine actually did.

git clone https://github.com/LinusInnovator/wtf.git

cd wtf

npm install

npm run build

npm test

npm run gauntletBuilt by @LinusInnovator. Explored in depth at Great Delights.