Harness Engineering: Anatomy of 11 Production Coding Agents
A source-code audit of 11 coding agent harnesses reveals shared architecture, 29 design patterns, and the shift from tool to platform in 2026.
A new empirical study maps the internals of eleven production coding agent harnesses — including Claude Code, Codex CLI, Gemini CLI, OpenHands, and Aider — to formalize "harness engineering," the discipline of designing the runtime that wraps an LLM with tools, context management, safety controls, and orchestration. Analyzing roughly four million lines of Python, TypeScript, and Rust, the authors define seven canonical subsystems shared across these systems and catalog 29 recurring design patterns alongside 13 cross-cutting observations.
Two notable gaps persist across the entire corpus: none of the surveyed runtimes import a general-purpose agentic framework, and none use vector embeddings for code retrieval — the field still relies on hand-rolled async loops and deterministic search. Skills slightly outpace MCP adoption, while ACP now ships in six systems, introducing a new role termed "harness hosting." A longitudinal diff across one quarter shows harnesses converging toward each other and shifting behavioral policy from prompt text into structured configuration.
The paper's central claim: in the first half of 2026, the coding harness evolved from a simple tool wrapper into a platform. It closes with 18 concrete design recommendations and a minimal 90-line reference implementation for engineers building their own agent runtimes.
This synthesis was produced from its source by AI; there is no human editor or manual review step. How we work