Harness Engineering with Claude Code
I built three harnesses, each iteration built on findings from the last: v1 exposed the spec as the real bottleneck, v2 added a separate evaluator agent since self-evaluation skews positive by design, v3 took over unfinished code and closed 47 features with 821 backend tests.