Best of Agent Harnesses and Harness Techniques

🏆  Curated list of AI agent harnesses, orchestration frameworks, and harness techniques for reliable agentic systems.

🌐 Browse the searchable site — one page per harness, filter by capability, autonomy & recovery.

🧰 Templates and Playbooks: copy-paste setup files and step-by-step guides for the harnesses in this list.

🤖 Agents can query this list — an MCP server (recommend, pick_harness, …), llms.txt & JSON, so your agent recommends harnesses too.

What is an agent harness?

A model answers; an agent acts. An agent harness is the runtime that turns one into the other: the model thinks, the harness decides what that thinking is allowed to touch.

Simon Willison's definition of the agent itself is the cleanest: "an LLM agent runs tools in a loop to achieve a goal." The harness is everything around that loop: which tools exist, what needs approval, what the model sees each turn, what survives a crash. Andrej Karpathy named the architecture back in 2023: the model is "the kernel process of a new Operating System", and the harness is the rest of that OS, its scheduler, permissions, and memory. The SWE-agent paper proved the stakes by coining the agent-computer interface: how tools and feedback are presented changes what a model can do, independent of the model. The field's advice has since converged on investing here rather than in framework plumbing, from Anthropic's build-simple guidance to Jerry Liu's argument that the framework era is over and the layers that matter now are skills, tools, and context quality. Those are the layers this list catalogs.

Better models make harnesses more important: more capabilities mean more failure modes, and production needs retry logic, fallbacks, and validation. Harness quality, not just model quality, determines whether agents actually ship. This list ranks projects by relevance to harness concerns (environment, orchestration, lifecycle, guardrails) and by stars/activity.

The benchmark data now backs this up. On SWE-bench Pro, "swapping the agent harness changed pass@1 more than many model upgrades do" (AINews, Aug 8 2026, citing analysis by @joelniklaus). Same model, different harness: 23% to 52% pass@1 on GLM-5.2, and 15% to 36% on Gemma 4 26B. Harness rankings barely transfer across models (rank correlation -0.05), so a small model in the right harness can approach a much larger model in the wrong one.

The gap is widest on long-horizon work. A bare frontier model was verified at about 30% on ARC-AGI-3; Prime Agent's harness took Opus 5 to 95.5%, and the YC Paper Club talk on why the harness matters more than the model (September 2026) walks through how. The gains come from a new class of self-improving harnesses, not from piling on scaffolding: the best of them stay thin and expose only what the model cannot do for itself. And because rankings barely transfer across models, the harness choice is a pairing that must be re-asked whenever the model changes. Who says so, what they measured, and what the claim does not mean is its own page.

That is the problem the MCP server in this repo solves. Point your agent at it and it can call recommend or pick_harness to choose a harness matched to your model and task, instead of inheriting whichever harness someone else benchmarked.

The landscape at a glance

Every project in the list, plotted by adoption surface area (the simplicity ↔ capability axis) against GitHub stars. Colors are categories; the largest projects in each tier are labeled.

The same projects placed by how much unsupervised rope they're designed to give (autonomy) and what happens when a run dies (recovery). In the tables below, ★ marks headless-ready projects and ✱ marks durable ones. Both charts regenerate from the list data on every refresh.

Start with the guide, then the head-to-head decision pages — grounded in the same data as the tables below:

- Why the harness matters more than the model: the measurements (ARC-AGI-3 30% to 95.5% on the same weights), who says so, and what the claim does not mean

- How to pick a harness: six questions that turn this list into a decision, plus the chart to internalize first (the harness moves scores more than the model)

- How to test-drive a harness: the two-week trial protocol, with a fair setup, tasks from your own repos, seven measurements, and the walk-away test

- The best AI agent harnesses in 2026, ranked by category: the top three per category from this week's data, regenerated on every refresh

- OpenClaw vs Hermes — the always-on personal-agent debate: presence vs discipline, plus what the field reports actually say

- Managed vs self-hosted always-on agents: Grok Bot, Claude Managed Agents, QM, OpenClaw, Hermes, and OpenJarvis; who owns the computer and who pays for idle time

- Terminal coding agents — opencode vs Codex vs Gemini CLI vs crush vs goose

- Multi-agent orchestration — OpenAI Agents SDK vs CrewAI vs AutoGen vs LangGraph

- Agent memory layers — Mem0 vs Letta vs claude-mem

- Agent sandboxing: what it is, the key concepts, and the field (E2B vs Daytona vs Modal and more)

- Agent evals (SWE-bench vs inspect_ai vs Terminal-Bench)

- Eval and observability platforms (Langfuse vs LangSmith vs Braintrust vs Phoenix)

- Browser agents (browser-use vs Stagehand vs Playwright MCP vs chrome-devtools-mcp)

- Browser infrastructure (Browserbase vs Steel vs Hyperbrowser)

- Claude Code skill packs (superpowers vs GStack vs get-shit-done vs Anthropic Skills)

- Context files for agents (AGENTS.md vs CLAUDE.md vs skills vs MCP tool search)

The list tells you which harness to use. These tell you how to set it up. Templates are files you copy into your project; Playbooks walk you through one task, step by step.

Templates

- One AGENTS.md for every coding agent: One briefing file that Codex, Claude Code, Cursor, OpenCode, GitHub Copilot, Gemini CLI, and Aider all read, so you write your build commands, conventions, and hard rules once instead of once per tool.

- Safe Claude Code settings: permissions and a guard hook: Project settings that stop Claude Code from force-pushing, wiping work, reading secrets, or piping downloads into a shell, while leaving everyday commands alone. Two files, copy and commit.

- Minimal agent harness in Python: A working coding-agent harness in about 180 lines of Python: the loop, three tools, a permission gate, a context file, a turn budget, and a transcript you can resume after a crash. Copy it to learn how harnesses work, or as the start of your own.

Playbooks

- Build your own agent harness: Build a working coding agent in about an hour: a loop, three tools, permissions, a context file, a budget, and crash recovery, in one Python file you fully understand. You finish with the minimal harness template running on your own repo.

- Write one AGENTS.md for every coding agent: Write one briefing file that Codex, Claude Code, Cursor, OpenCode, Copilot, Gemini CLI, and Aider all read, test that each tool actually loaded it, and keep it short enough to help instead of hurt. You finish with the AGENTS.md template filled in for your repo.

Reader's index: pick by what you want to do, not by category. Tag chips (e.g. mcp · memory) next to each row let you cross-filter by capability — see TAGS.md for the full cross-reference.

- I want a turnkey coding agent today — opencode, Cline, Codex, Gemini CLI, OpenHands, crush, Prime Agent · see Coding agent products (IDEs, CLIs, full suites)

- I want an always-on personal agent that lives in my chat apps — OpenClaw, Hermes, Khoj, Agent Zero, OpenHarness (HKUDS), QM, OpenJarvis · see Personal agent runtimes

- I want to extend Claude Code, Codex, or OpenCode with skills and slash commands — Anthropic Skills, wshobson/agents, superpowers, GStack, pmstack · see Coding harness configs and SDKs

- I want to build my own coding harness from scratch — Claude Agent SDK, Google ADK, AutoHarness, SWE-agent, RepoMaster, claw-code-agent · see Coding harness configs and SDKs

- I want a drop-in memory layer for agents — Mem0, Graphiti (Zep), claude-mem, agentlog, letta · see Plugins, MCPs, CLI tools

- I want to plug hundreds to thousands of tools without context bloat — MCP-Zero, ToolGen, ToolRAG, langgraph-bigtool · see Progressive disclosure harnesses

- I want multi-agent orchestration — openai-agents-python, crewAI, autogen, Microsoft Agent Framework, PraisonAI, agent-squad · see Multi-agent and orchestration

- I want a general LLM app framework — langgraph, langchain, llama-index, pydantic-ai, agno · see Frameworks

- I want low-code / visual workflows — langflow, Flowise, Dify, n8n · see Frameworks

- I want browser-using agents — browser-use, Stagehand, WebVoyager, puppeteer-real-browser-mcp · see Plugins, MCPs, CLI tools

- I want sandboxed code execution for agent-generated code — E2B, Agent Sandbox, Daytona, smolagents, OpenHands · see Libraries and SDKs

- I want to evaluate or benchmark agents — SWE-bench, Terminal-Bench, AgencyBench, inspect_ai, WebArena, VitaBench · see Evaluation and benchmarking harnesses

- I want a deep research / autonomous research agent — deepagents, gpt-researcher, openagents · see Research and task-specific harnesses

- I want a provider-agnostic LLM pipe (not a framework) — LiteLLM, vercel/ai · see Libraries and SDKs

This list is also published in machine-readable form, so coding agents and research agents can recommend harnesses — not just humans browsing GitHub:

- harnesses.json — every project with category, complexity tier, capability tags, stars, license signal, and a concrete example link, plus the full use-case index.

- llms.txt — the entire list in one agent-readable file. Point any agent at the raw URL.

- MCP server — recommend(one opinionated pick + alternatives + what to avoid, e.g. repos flagged for star manipulation),compare/compare_for(2–4 harnesses side by side — by id or by task — who leads on which axis incl. researched sandboxing/memory/hooks/prompt-optimization ratings, graveyard warnings, the matching decision guide),pick_harness(ranked, with complexity/autonomy/recovery filters),pick_infrastructure(picks at any level of the infra stack plus a live GitHub/Hacker News discovery pass, so answers aren't limited to this list),search_harnesses,get_harness,list_categories, pluslist_comparisons/get_comparisonfor the decision guides andlist_templates/get_template/list_playbooks/get_playbookso your agent can install a template for you. Published to PyPI and the official MCP registry asio.github.RyanAlberts/agent-harnesses. One-line install (needs uv):

claude mcp add agent-harnesses -- uvx agent-harnesses-mcp

Don't just read the list — agents/ ships three agent skeletons: open-source agents that run on the AI subscription you already pay for. Clone the file, customize the instructions, done. All three work against the current week's data and deliver to Slack or Notion when either is connected:

- harness-scout — describe what you're building; it picks your harness, with evidence and a graveyard check.

- stack-auditor — flags the harnesses in your codebase that died, and can trace your agent session logs to show how the harness steers your technical decisions.

- harness-radar — weekly movement briefing: climbers, arrivals, deaths, graduations.

curl -fsSL https://raw.githubusercontent.com/RyanAlberts/best-of-Agent-Harnesses/main/agents/harness-scout.md -o .claude/agents/harness-scout.md

- ⭐ Stars — GitHub star count, captured 2026-09-27; tables sort by stars descending.

- ⚖️ Simplicity ↔ capability — adoption surface, 4 tiers: super simple (a format, one concept) → mostly simple (thin layer) → slightly complex (real SDK) → complex (product suite).

- ★ Headless-ready — designed for unattended runs, batches, and fleets (the top of the autonomy scale: step-gated → checkpoint-gated → bounded → headless).

- ✱ Durable — persisted execution state survives restarts mid-task (the top of the recovery scale: none → retry → resumable → durable).

- ✅ Open source — ✅ standard OSS license · ⚠️ source-available/restricted · ❓ no or unclear license.

- 🏷️ Tags — capability chips auto-derived from descriptions; full cross-reference in TAGS.md.

- 🎯 Examples — one concrete "show me it in action" link per project, not a docs root.

Every project's full autonomy and recovery tier is plotted in the grid above and carried in harnesses.json and llms.txt; scores are editorial, from public docs — maintainer corrections via issue/PR are merged fast.

Progressive disclosure harnesses

Formats, runtimes, and patterns that reveal context, tools, or instructions in layers—index first, details on demand—to control tokens and improve agent focus (the "map, not encyclopedia" principle).

Coding agent products (IDEs, CLIs, full suites)

Turnkey coding agents you install and run: IDE extensions, terminal CLIs, Dockerized workspaces. Each entry notes which part is the harness (the agent loop, tool wiring, approval model) versus the UI shell (VS Code extension, TUI, browser client).

Coding harness configs and SDKs

Skill packs, slash-command libraries, meta-prompting frameworks, and official SDKs that give you the harness (the agent loop, planning, memory, hooks) without bundling a specific IDE or CLI shell.

Always-on, self-hosted agents you run as a daemon and talk to from chat apps: gateway runtimes, second brains, and self-improving assistants. The agent as a product you operate, not a library you build with.

General-purpose agent and LLM application frameworks (the app layer, not harnesses per se).

Multi-agent and orchestration

Harnesses and patterns for multi-agent coordination and handoffs.

IDE plugins, concrete MCP servers, and CLI tools that give agents tools and context.

Persistent memory layers that give agents recall across turns and sessions: knowledge graphs, vector stores, and session-capture tools that survive a restart. The state a harness needs but rarely ships with.

Evaluation and benchmarking harnesses

Agentic eval systems, reasoning benchmarks, and open agent benchmarks.

Observability and eval-ops

Tracing, monitoring, and production evaluation for live agent runs: capture every step, tool call, and token, then score and debug in the loop. Distinct from the fixed-task benchmarks above—this is what you run against your own traffic.

Research and task-specific harnesses

Deep research, document QA, and domain-specific agent loops.

Lightweight runtimes, tool loops, and provider-agnostic harness primitives.

Archived upstream, or flagged for curation integrity (e.g. suspected star manipulation). Kept here — not deleted — for citation and transparency; excluded from the ranked count, the landscape chart, and harnesses.json's main list. Curation is the point: a starred repo is not automatically a credible one.

Up-and-coming candidates — surfaced by the weekly discovery scan or submitted by the community — that haven't cleared the curation bar or a vetting pass yet. Stars refresh weekly from the discovery queue; descriptions are the projects' own, unvetted. Entries graduate into the ranked list above or drop off.

Which agent harnesses can run unattended (headless)?

Harnesses designed for unattended runs, batches, and fleets: opencode, OpenHands, goose, Symphony, Prime Agent, SWE-agent, Claude Agent SDK, RepoMaster.

Which agent harnesses survive a crash mid-task (durable)?

Harnesses whose execution state persists across restarts: langgraph-bigtool, QM, n8n, langgraph, mastra, letta, deepagents, pydantic-ai.

How many of these agent harnesses are open source?

124 of 167 carry a standard open-source license; the rest are source-available or unclear, and flagged per row.

What is an agent harness?

The runtime that turns a model into an agent: it decides what the model's reasoning is allowed to touch, and supplies the orchestration, tool wiring, memory, error recovery, and guardrails around per-turn inference.

Does the harness matter more than the model?

Often, yes, and measurably: the same weights scored about 30% on ARC-AGI-3 as a bare model and 95.5% inside the Prime Agent harness, and on SWE-bench Pro swapping only the harness moved GLM-5.2 from 23% to 52%. Harness rankings barely transfer across models (rank correlation about -0.05), so pick the harness and the model as a pair, and re-pick when the model changes. Sources and the full argument

Is Grok Bot an agent harness?

Yes, a managed one: xAI owns the loop, the tool wiring, the memory, and the approval rules, and every Bot on an account shares one cloud computer. It is not in the ranked list because the list ranks open repositories; the managed-vs-self-hosted guide compares it with Claude Managed Agents, QM, OpenClaw, Hermes, and OpenJarvis. Sources and the full argument

By relevance to harness concerns (environment, orchestration, lifecycle, guardrails) and by GitHub stars (captured 2026-09-27); each project also carries an adoption-surface tier and autonomy/recovery scores.

How can an AI agent use this list directly?

Three machine-readable surfaces: harnesses.json (structured), llms.txt (one file), and an MCP server (uvx agent-harnesses-mcp) exposing recommend, compare, pick_harness, and search_harnesses.

- Awesome: Awesome lists on many topics

- OpenAI – Harness engineering: Environment design, intent, feedback loops, repo-as-system-of-record

- Anthropic – Effective harnesses for long-running agents: Session bridging, feature lists, incremental progress, testing

- Aakash Gupta (Medium) – 2026 is agent harnesses: Harness as moat, minimal intervention, progressive disclosure

- YC Paper Club (video) – Why the harness matters more than the model: September 2026 session with the authors of Prime Agent, OpenJarvis, and QM; the history of harnesses runs from 7:00 to 17:00

- LangChain, Anthropic, OpenAI: Official docs for major agent platforms

🧡 Thank you, contributors

The people who stopped mid-scroll, found a gap, and wrote it up — this list is better for each of them:

@baskduf — harness-starter-kit · @ahwurm — LocalHarness · @liviux — LoopTroop · @rishabhpoddar — TeamCopilot, on the radar · @abilliontokens — oh-my-pi · @pranshuchittora — agent-qa · @madarco — AgentBox · @jmthomasofficial — JMT x402 Agent Tools · @ShukantPal — Proliferate · @hjqcan — GoodMemory · @777genius — Agent Teams AI · @S1LV3RJ1NX — mcp-guardian · @AmariahAK — Atlarix, Update Atlarix entry · @hardness1020 — awesome-agent-architecture · @razzant — Claudexor, Ouroboros · @allenshi16 — Nexus AI Pulse, HomeOffice AI · @ryanpettry — fractal · @reacher-z — ClawBench evaluation harness · @HelpMatey — Pin mcp2 · @rxdt — L∞pGate · @msaleme — red-team-blue-team-agent-fabric · @JackChen-me — Open Multi-Agent · @Q00 — Ouroboros · @danawoodman — PR #78 · @JanYork — LWC Local Wiki CLI · @labmimors — MCP Lens · @elmariachi111 — Prime Agent · @denial123789 — SandBase Harness · @vshulcz — deja · @hasmcp-dev — MCP spec conformance · @kuanzema — HarnessRouter · @liyangbing — SandBase Harness · @BlueSkyID666 — Orkas · @jackispm — Nausicaa · @xizhuomengcontin — OrcaReplay · @InsightFactoryAPP — YYLO, YYLO Benchmark · @AbdulDavids — Gram · @scgopi — GraphCode · @denggui-ai — Minimal Harness · @JakeSelby — Agent Harness · @OlyaTi — Mnemoverse · @IRONICBo — Jev Social

Accepted submissions land with co-author credit on the commit that ships them. Promising projects that are still early aren't turned away — they get pinned to 🔭 On the radar and graduate as they grow. Add yours →

Contributions are welcome. To add or suggest projects:

- Open an issue with the repo URL, category, and a short description.

- Or submit a pull request against scripts/generate.py — this README, projects.yaml, and TAGS.md are generated from it, so direct edits to them can't merge.

Promising projects that don't clear the curation bar yet get pinned to 🔭 On the radar — a submission that lands there isn't rejected, it's queued.

For contribution guidelines, see CONTRIBUTING.md and the Code of Conduct.

If your project is in this list, you're welcome to show it in your README:

[![Best of Agent Harnesses](https://img.shields.io/badge/%F0%9F%8F%86_Best_of-Agent_Harnesses-5ac4bf)](https://github.com/RyanAlberts/best-of-Agent-Harnesses)