npx skills add ...
npx skills add affaan-m/ecc --skill gan-style-harness
GAN-inspired Generator-Evaluator agent harness for building high-quality applications autonomously. Based on Anthropic's March 2026 harness design paper. Use when a feature should be built autonomously through generator and evaluator iteration until it clears a quality bar.
npx skills add affaan-m/ecc --skill gan-style-harness
Inspired by Anthropic's Harness Design for Long-Running Application Development (March 24, 2026)
A multi-agent harness that separates generation from evaluation, creating an adversarial feedback loop that drives quality far beyond what a single agent can achieve.
When asked to evaluate their own work, agents are pathological optimists — they praise mediocre output and talk themselves out of legitimate issues. But engineering a separate evaluator to be ruthlessly strict is far more tractable than teaching a generator to self-critique.
This is the same dynamic as GANs (Generative Adversarial Networks): the Generator produces, the Evaluator critiques, and that feedback drives the next iteration.
claude -p)Role: Product manager — expands a brief prompt into a full product specification.
Key behaviors:
Model: Sonnet by default; raise via GAN_PLANNER_MODEL=opus for deeper spec expansion
Role: Developer — implements features according to the spec.
Key behaviors:
Model: Sonnet by default; raise via GAN_GENERATOR_MODEL=opus for maximum coding capability
Role: QA engineer — tests the live running application, not just code.
Key behaviors:
Model: Sonnet by default; raise via GAN_EVALUATOR_MODEL=opus for stronger judgment + tool use
The default four criteria, each scored 1-10:
The harness should simplify as models improve. Following Anthropic's evolution:
Key principle: Every harness component encodes an assumption about what the model can't do alone. When models improve, re-test those assumptions. Strip away what's no longer needed.
| Variable | Default | Description |
|---|---|---|
GAN_MAX_ITERATIONS | 15 | Maximum generator-evaluator cycles |
GAN_PASS_THRESHOLD | 7.0 | Weighted score to pass (1-10) |
GAN_PLANNER_MODEL | sonnet | Model for planning agent |
GAN_GENERATOR_MODEL | sonnet | Model for generator agent |
GAN_EVALUATOR_MODEL | sonnet | Model for evaluator agent |
GAN_EVAL_CRITERIA | design,originality,craft,functionality | Comma-separated criteria |
GAN_DEV_SERVER_PORT | 3000 | Port for the live app |
GAN_DEV_SERVER_CMD | npm run dev | Command to start dev server |
GAN_PROJECT_DIR | . | Project working directory |
GAN_SKIP_PLANNER | false | Skip planner, use spec directly |
GAN_EVAL_MODE | playwright | playwright, screenshot, or code-only |
| Mode | Tools | Best For |
|---|---|---|
playwright | Browser MCP + live interaction | Full-stack apps with UI |
screenshot | Screenshot + visual analysis | Static sites, design-only |
code-only | Tests + linting + build | APIs, libraries, CLI tools |
Evaluator too lenient — If the evaluator passes everything on iteration 1, your rubric is too generous. Tighten scoring criteria and add explicit penalties for common AI patterns.
Generator ignoring feedback — Ensure feedback is passed as a file, not inline. The generator should read feedback-NNN.md at the start of each iteration.
Infinite loops — Always set GAN_MAX_ITERATIONS. If the generator can't improve past a score plateau after 3 iterations, stop and flag for human review.
Evaluator testing superficially — The evaluator must use Playwright to interact with the live app, not just screenshot it. Click buttons, fill forms, test error states.
Evaluator praising its own fixes — Never let the evaluator suggest fixes and then evaluate those fixes. The evaluator only critiques; the generator fixes.
Context exhaustion — For long sessions, use Claude Agent SDK's automatic compaction or reset context between major phases.
Based on Anthropic's published results:
| Metric | Solo Agent | GAN Harness | Improvement |
|---|---|---|---|
| Time | 20 min | 4-6 hours | 12-18x longer |
| Cost | $9 | $125-200 | 14-22x more |
| Quality | Barely functional | Production-ready | Phase change |
| Core features | Broken | All working | N/A |
| Design | Generic AI slop | Distinctive, polished | N/A |
The tradeoff is clear: ~20x more time and cost for a qualitative leap in output quality. This is for projects where quality matters.