A deterministic, multi-agent methodology for AI-assisted software development.

Several coding agents working on one repository fail in predictable ways: they collide on branches and on shared context files, claim results they never measured, and write tests that pass whatever the code does. AFAW is a repository-level method and boilerplate that lets agents write code in parallel while deterministic tools decide whether the result is acceptable and a human approves every merge.

AFAW is the method; the code here is one way to implement it. See what is core and what you can replace.

White paper (PDF): docs/paper/afaw.pdf · LaTeX source: docs/paper/afaw.tex · DOI 10.5281/zenodo.23050310

Live demo: agustindiazcano.github.io/afaw-anti-fragile-agentic-workflow · this repository's own dashboard: /live/

A dashboard built by scripts/build_state.py from mock data: a fictional project whose every task and number is invented (examples/mock_dashboard.py). It shows each traffic-light rule firing: CI failing, a blocker not done, no activity for three days, nothing measured.

Build it locally: python -m examples.mock_dashboard mock-output and open mock-output/index.html.

-

It costs less. Without it, an agent runs the whole test suite on your machine, the machine is slow, tests fail for reasons that have nothing to do with the code, and the agent retries for an hour. Every retry sends its whole context again: that is where the token bill goes. With AFAW the agent runs only the test it is writing; the cloud runs the rest in parallel and answers once. A few minutes of CI cost far less than an hour of an agent going in circles.

-

It works alone and in a team. Alone, you let the agents run and nobody checks them: they commit to the wrong branch, write tests that pass no matter what, and report results they never measured. The checks, the branch protection and the hooks do that reviewing for you.

-

With 2 or 10 people, everyone sees what the agents did. Each person runs their own agents, and without a shared record nobody knows what the others' agents changed, why, or where each task stands. Here every task, decision and status is in the same place and in the same format, for every person and every agent: the board, the generated LASTCONTEXT.mdand the decision records. You set it up once, from the template.

-

The project documents itself. Every task leaves what was done, what is left and, for each decision, why it was taken. When agents do the work, that is usually lost: afterwards nobody knows what was done or why.

-

A new agent understands the project in minutes. Instead of reading every file from scratch, it reads one generated index ( LASTCONTEXT.md) and follows links to the decisions it needs. Less reading means fewer tokens and less time, and the saving grows with the project.

-

Fewer status meetings. "What did you do, where does it stand, what's blocked, why was this decided": the board and the records already answer it, and an optional assistant can answer it in a chat. See Ask the project.

These are the expected gains, not measured ones yet. The dashboard already records timings and CI cost, and the paper's evaluation protocol says how to measure all four.

The method does not require these exact tools. Write your tests however you like, use any CI, skip mutation testing outside critical code. Some parts, though, are what make it work: remove them and the problem they solve comes back.

Because everything is recorded as the work happens, you can put an assistant with RAG on top: it searches the project's records and answers questions in a chat. It is not part of the method, just an optional use of what the method already records.

What it replaces: status meetings. "What did you do, where does it stand, what's blocked, why was this decided." The board, LASTCONTEXT.md and the decision records already answer that. The assistant answers the same questions any time, to anyone, without taking an hour from five people.

Why it works well here. Every document is about one thing: one decision per ADR, one trap per gotcha, one delta per task. They split into clean pieces. The front matter says whether a decision is still in force, so the assistant can skip the ones that were replaced. LASTCONTEXT.md is the entry point. And every fact has a path, so each answer can cite where it came from and anyone can check it.

What it doesn't replace.

- Meetings where things get decided. Priorities, trade-offs, disagreements still need people. The records say what was decided, not what to decide. Those meetings do get shorter: nobody spends the first half working out where things stand.

- What nobody wrote down. A decision made in a hallway and never turned into an ADR doesn't exist for the assistant.

- Records that went stale. If an ADR stopped being true and nobody replaced it, the assistant answers with confidence and gets it wrong.

Two rules for the assistant.

- Numbers are copied, not paraphrased. Progress, lead time, CI cost: quote them from the derived views. An assistant that "summarizes" metrics can invent them.

- Read-only. It answers; it never writes to the repository, opens pull requests or changes state.

- Demo dashboard

- Why use it

- Method and this implementation

- Optional: ask the project

- Core principle

- Failure modes and mechanisms

- System overview

- How it works

- The life of a task

- Directory structure

- Getting started

- Daily workflow

- Scope and limits

- Rebuilding the diagrams and the paper

- Citation, author, license

A model proposes and writes code; its own assessment is never trusted. Every number (tests, types, coverage, mutation score, task timings, the traffic light) comes from a deterministic tool, and a value an agent could have made up is rejected where it would be declared.

Colors in all diagrams: blue = agent action, amber = human, green = CI or deterministic check, purple = state files, red = blocked.

An agent writes only two files: its task file state/tasks/task_NNN.json (declared values: title, role, type, status, owner, priority, difficulty, dependencies, branch) and its delta context/tasks/task_NNN_context.json (summary, next, files_touched). Their shape is defined once, in state/schemas/.

Everything shared is derived, never committed: python -m scripts.build_state --ref origin/main builds the pending view, the history of done tasks, the metrics, the dashboard and the LASTCONTEXT.md files into build/. Timestamps and CI results are measured from git and the CI API (python -m scripts.collect_facts), never declared; the traffic light is a rule:

- 🔴 a task it depends on is not done, or CI fails on the latest commit of its branch

- 🟡 in progress with no commit or CI run for more than stale_days

- 🟢 otherwise; "no data" when nothing was measured

State says what; prose says why. Decisions are ADRs in docs/adr/, operational traps are gotchas in docs/gotchas/ (templates in docs/templates/). The generated LASTCONTEXT.md indexes them with one line each, next to the current state, what waits on the human and the next tasks. A record leaves the index only for an explicit reason (superseded, deprecated, resolved, or promoted to a mechanism), never for its age, and a budget tells the human when to prune. See ADR 0001.

Locally, the agent runs only the test it is writing. CI runs everything else, including the red-first check (python -m scripts.red_check): each test a pull request adds must fail on the code before the change.

Only new or modified files are mutated. In critical modules a file below the minimum blocks the pull request. The gate is K / (N − E_human): only the human records an equivalent mutant, in state/equivalent_mutants.json, bound to the file's hash (python -m scripts.mutation_gate). The mutation engine itself is project-specific.

No commits or pushes to main, no merge without an explicit human directive, and the human approves every pull request. Rules in AGENTS.md are advice, so they are also enforced by hooks, CI and branch protection.

A static HTML page (no scripts) with progress, tasks by status, tasks merged per day, the board with ages and lights, timing (agent time, review latency, lead time) and CI cost (runs, queue time, job minutes). .github/workflows/dashboard.yml rebuilds it on every push to main, every hour and on demand, and publishes it on GitHub Pages: the demo at the site root and this repository's own dashboard at /live/. A Pages site is public: do not publish the dashboard of a private project.

.

├── AGENTS.md # The rules: single source for every agent

├── CLAUDE.md # Imports AGENTS.md

├── PENDING.md # The human's roadmap (hook: never deleted or emptied)

├── .claude/

│ ├── settings.json # Hook wiring

│ ├── hooks/ # session_start (builds and injects LASTCONTEXT.md),

│ │ # safety_guard (deny/ask), lint_check

│ └── skills/ # commit, ship, tests, push-dev, trash

├── .github/

│ ├── pull_request_template.md

│ └── workflows/

│ ├── checks.yml # Every PR: task files, delta, red-first, docs, views

│ ├── scripts-ci.yml # Lint, types, tests (skipped on documents-only PRs)

│ └── dashboard.yml # Builds and publishes the dashboard

├── state/

│ ├── config.json # Code folders, mutation settings, dashboard settings

│ ├── schemas/ # JSON Schemas: task, delta, config, equivalent mutants

│ └── tasks/task_NNN.json # One file per task (never deleted)

├── context/tasks/ # One delta per task

├── docs/

│ ├── adr/ # Architecture decision records

│ ├── gotchas/ # Operational traps, closed as resolved or promoted

│ ├── templates/ # ADR and gotcha templates

│ ├── diagrams/ and img/ # Graphviz sources and rendered figures

│ ├── screenshots/ # Screenshots of the demo dashboard

│ └── paper/ # The white paper: afaw.pdf and its source afaw.tex

├── scripts/

│ ├── afaw_state/ # Model, facts, lights, durations, views, dashboard

│ ├── build_state.py # Derived views from a git ref

│ ├── collect_facts.py # Measured facts from git and the GitHub API

│ ├── check_task_files.py # Lifecycle, declared values only, unique ids

│ ├── check_task_delta.py # files_touched equals the diff (--write fills it)

│ ├── check_docs.py # ADRs, gotchas, links

│ ├── red_check.py # Added tests must fail on the old code

│ ├── mutation_targets.py # Mandatory and optional files to mutate

│ ├── mutation_gate.py # Score with human-accepted equivalences only

│ ├── rules_to_tex.py # The paper's appendix from AGENTS.md

│ └── setup_protection.sh # Branch protection and required checks

├── examples/mock_dashboard.py # The demo: dashboard and LASTCONTEXT.md from mock data

├── src/ # The project's code

└── tests/tooling/ # Tests of the scripts and hooks

build/ (the derived views) is ignored by git.

GitHub does not copy branch protection, required checks or Pages settings to a repository created from a template.

- Use the template: GitHub → Use this template.

- Configure state/config.json: code folders,mutation_critical_paths,mutation_min_score, and thedashboardsettings.

- Protect main:bash scripts/setup_protection.sh(needs theghCLI) requireslint,typecheck,testsandstate-and-docs, and blocks direct pushes.

- Dashboard (public repositories): Settings → Pages → Source: GitHub Actions.

- Verify: open a test pull request and check that the workflows run.

- Start: open a terminal per agent and answer the ROLE:question.

- One terminal per agent; answer its ROLE:. The session-start hook injects the generatedLASTCONTEXT.mdandPENDING.md.

- The agent creates its branch, works with TDD running only its own test, and pushes; CI verifies the rest.

- At the end it writes summaryandnext, fillsfiles_touchedwithpython -m scripts.check_task_delta --base origin/main --write task_NNN, records any decision as an ADR, and asks before opening the pull request.

- You review and merge. Nothing else to approve: the views are derived.

- Ask for status at any time, or open the dashboard.

- The reference rules assume a Python backend (async, pytest,mypy --strict, Postgres, Terraform). Adapt sections 4–10 and 14–16 ofAGENTS.mdfor another stack.

- AFAW has not been validated in a controlled study and makes no claim about speed, cost or defect rates. The paper proposes an evaluation protocol; the dashboard's timings and CI cost are the raw material.

- Mutation score measures how sensitive the tests are to changes, not whether the code meets its requirements.

- The red-first check shows that a new test fails on the old code, not that it fails for the right reason.

- Checks verify the ADRs that exist; a decision that was never recorded is left to the reviewer.

- Documentation that stops being true misleads more than none. Derived views are rebuilt and decisions can be superseded, but nothing detects a decision record whose text went stale while nobody replaced it.

for f in docs/diagrams/*.dot; do n=$(basename "$f" .dot); dot -Tsvg "$f" -o "docs/img/$n.svg"; dot -Tpng -Gdpi=200 "$f" -o "docs/img/$n.png"; doneThe paper: see docs/paper/README.md. A pull request that changes afaw.tex, AGENTS.md or docs/img/ must also commit the rebuilt docs/paper/afaw.pdf; CI checks it.

Agustin Diaz-Cano, MS Candidate, Information Systems Engineering (UTN). ORCID 0009-0001-4336-490X.

MIT