better-harness logo

better-harness

Help your coding agents (Claude Code, Codex, Qoder, Cursor, and other coding agents) get better at getting better.

node >=22.20.0
Website GitHub

What is it?

What it is

A tool that reviews AI coding workflows by analyzing five dimensions of the Agent Work Loop (Task Understanding, Controlled Execution, Change Validation, Reliable Delivery, Learning Capture) to identify gaps and provide actionable findings based on evidence.

Why it exists

To address weak points in AI coding workflows by making findings explicit and tied to evidence, helping teams improve their processes through actionable insights.

Who should use it

Software engineers using AI coding agentsDevOps engineers looking to improve AI delivery safeguardsEngineering managers monitoring AI coding workflow qualityTeams using Claude Code, Cursor, GitHub Copilot, or QoderAI agent developers seeking workflow evaluation metrics

Who should avoid it

Users not using supported AI coding agentsDevelopers looking for a simple code completion tool rather than a workflow reviewerUsers without Node.js environment for local development

How it works

A quick walkthrough in plain English

How better-harness works

Step 1 of 3

You interact with it

Open better-harness, send a request, or connect it to your stack.

Features

Reviews AI coding workflow across multiple agents (Claude, Codex, Cursor, etc.)
Evaluates five dimensions of Agent Work Loop (Task Understanding, Execution, Validation, Delivery, Learning)
Generates evidence-bounded reports with prioritized findings
Open-source with customizable engineering practices and evaluation models
Integrates with Qoder, Claude Code, GitHub Copilot, and other coding agents
Provides transparent evidence tracking instead of inferred scores

Advantages

  • Transparency in workflow analysis through visible evidence
  • Structured evaluation of critical agent workflow dimensions
  • Multi-agent support with standardized reporting
  • Open-source flexibility for customization and extension
  • Community-driven improvements via contribution model
  • Focus on systemic improvements over superficial fixes

Disadvantages

  • Requires specific coding agent integrations for full functionality
  • Setup complexity due to multiple host dependencies
  • Learning curve for understanding Agent Work Loop framework
  • Evidence coverage may vary between agents
  • May require additional tooling for optimal use

Installation

FAQ

What does Better Harness actually do?

Better Harness reviews your AI coding workflow by evaluating how agents understand tasks, execute changes, validate results, deliver code, and learn. It identifies gaps in these processes and turns them into prioritized findings with evidence-backed repair plans.

Which AI coding agents are supported?

It supports a wide range of hosts including Claude Code, Codex (Desktop and CLI), Qoder (Desktop and CLI), Cursor, Qwen Code, GitHub Copilot CLI, and Pi.

How do I run a review in my coding agent?

Once the plugin is installed, you can trigger a review by using the slash command: `/better-harness review this project's AI coding workflow and generate a report`.

What kind of output does the tool provide?

Depending on the host, it produces a self-contained HTML report, a Markdown report, or a Canvas report (for Qoder). These reports include prioritized findings, impact assessments, and scoped repair actions.

Does the tool infer behavior if evidence is missing?

No. Better Harness is designed to be deliberately honest; if behavior cannot be observed through session transcripts or project mechanisms, the report explicitly marks it as missing or partial rather than making unsupported claims.

Loading documentation…
View on GitHub

Featured in Videos

YouTube tutorials and walkthroughs for better-harness

Alternatives

Similar projects ranked by category, topics, and text overlap.

Compare
better-harness | MushyBook