The term Harness Engineering has established itself in recent months. The formula behind it is simple: Agent = Model + Harness. The harness is everything that makes up an AI agent, except the model itself. That includes instruction files like CLAUDE.md or AGENTS.md, skills, tools, tests, linters, hooks and CI gates.

The core point of Harness Engineering is a shift of focus. We stop telling the agent in the prompt what to do, and we start designing the system around the agent so that it cannot do otherwise. A sentence like “follow our architecture” in the prompt is a hope. A test that breaks the build is a guarantee.

In the AI Unified Process, the harness is one of three layers. In this post I use the aiup-petclinic project to show how the layers work together, why the CLAUDE.md there is deliberately small, why the architecture documents do the real work, and where tests stop and hooks begin.

The Three Layers of the AI Unified Process

The AI Unified Process distinguishes three layers:

What: What should the system do? This is where the requirements, the use case diagram, the use case specifications, the business rules and the entity model live. These artifacts are technology-neutral. They describe behavior from the point of view of the users and the business.

Harness: How do we make sure the agent builds the right thing the right way? This is where the instruction files, the skills, the architecture documents, the test strategy, the hooks and the CI gates live.

How: The actual code, the migrations, the tests. This is the output of the agent.

The order matters. Without What, the agent does not know what to build. Without Harness, it builds it in a way that works but does not fit our system.

Bigger Instruction Files Are Not Better

The common assumption is: the more context the agent has, the better it works. So everything ends up in the CLAUDE.md: project description, directory tree, architecture, coding conventions, test rules, commands.

A study from ETH Zurich from February 2026 (Gloaguen, Mündler, Müller, Raychev, Vechev: “Evaluating AGENTS.md”) tested this empirically. The results are sobering. Automatically generated context files lowered the success rate of the agents by 3 percent on average and raised the cost by over 20 percent. Human-written files brought only about 4 percent, with a similar increase in cost. And the agents did follow the instructions. The problem was not a lack of obedience, but that the instructions added noise, redundant steps and unnecessary constraints.

One detail of the study is interesting: when the researchers removed the existing documentation from the repositories, the context files suddenly helped. The conclusion: a context file is useful when it provides knowledge the agent cannot find anywhere else. It is useless when it repeats what is already in the documentation, and it is harmful when it fills the context with ballast.

For Harness Engineering this means: the CLAUDE.md is not a knowledge base. It is a map.

The CLAUDE.md in the PetClinic Project

The CLAUDE.md in aiup-petclinic is about 170 lines long and contains essentially four things:

  1. What the source of truth is. docs/ is the source of truth, not the code. If a use case and the code disagree, the use case wins. And which sensors check this claim.
  2. The stack and the commands. Java 25, Spring Boot 4.1, Vaadin 25.2, jOOQ 3.21, Flyway, PostgreSQL. How to build, test and regenerate the jOOQ classes. This is knowledge the agent cannot find anywhere else.
  3. When to read which document. Before implementing a use case: the use case specification. Before any change in src/main/java: docs/architecture/development.md. Before any test: docs/architecture/testing.md. Before touching .claude/: the section on agent guardrails.
  4. Which skills exist. The skills of the AI Unified Process are listed so the agent prefers them over ad-hoc generation.

What is not in the CLAUDE.md: the architecture rules themselves. No package layout, no jOOQ patterns, no Vaadin conventions. Those are in the architecture documents, and the agent reads them only when it needs them.

This is the difference between context that is always loaded and context that is loaded on demand. The CLAUDE.md is loaded into the context in every session. The architecture document only when the agent writes code. When writing a use case specification, it does not burden the context.

This has to be defended actively. When the traceability sensors were added (more on those later), the CLAUDE.md grew by a paragraph explaining what each sensor checks. The paragraph was correct, but it repeated what testing.md says in more detail. In the next commit it was cut down to a few sentences and a link: there are three sensors, the Status: line is therefore an assertion and not a label, and what each sensor checks is in testing.md. The map shows where the knowledge is. It is not the knowledge.

The Software Architecture Document as the Link

In the Rational Unified Process, the Software Architecture Document was the central document of the Elaboration phase. Many teams have dropped it in recent years. It was too much effort, nobody read it, it was out of date after three months.

With AI agents the math changes. The agent reads the document, on every task that touches code. And it follows it, as long as the document is short and concrete.

In the PetClinic project, the architecture lives in docs/architecture/, written as the 4+1 views Philippe Kruchten described in 1995. The use case view is the “+1” and sits outside the folder, because it is the base the architecture is built on. The other four views each answer one question:

ViewQuestionDocument
LogicalWhich business building blocks exist?logical.md
ProcessHow does the system behave at runtime?process.md
DevelopmentHow is the code organized, built and kept honest?development.md and testing.md
PhysicalWhere does the system run?physical.md

A README.md says which document answers which question, and eleven ADRs record the decisions behind the views. Three rules apply to all of these documents: text, not pictures, so diagrams are Mermaid or PlantUML; in the repository, not beside it, so architecture and code change in the same pull request; and short, not complete, because what is not written gets invented, and what is written at length gets skimmed.

For the agent, the most important view is development.md. It answers exactly the questions the agent would otherwise answer with what it has seen in training:

  • Package structure: package-by-feature under ai.unifiedprocess.petclinic. Each feature (owner, pet, visit, vet, welcome) has the sub-packages ui and domain. No separate service layer, no DTO layer, unless a use case demands it.
  • Data access: jOOQ, no JPA, no Spring Data repositories. Records are mapped with Records.mapping(Type::new), never with fetchInto. Nested records use row(...).mapping(Nested::new), parent-child relationships use multiset to avoid N+1.
  • Persistence stereotype: classes are named <Entity>Repository, annotated with @Repository so Spring’s exception translation applies, and they are the transaction boundary. A view is never transactional.
  • Vaadin conventions: one view per use case. Every route renders inside MainLayout. Styling only via LumoUtility, never via getStyle().set(). Validation in the form, not in the domain record.
  • Cross-feature rule: access to another feature only through its domain package, with one clearly defined exception for routing tokens and route parameter holders.

Without this document, the agent would probably have built OwnerController, OwnerService, OwnerRepository and OwnerDTO. Technically correct, but a different architecture, and after twenty use cases a project nobody understands anymore.

The business rules got a document of their own. A rule that belongs to a single use case stays in that use case as BR-NNN. A rule that several use cases share lives once in docs/business_rules.md as GR-NNN, and each use case references it instead of restating it. The reason is the same as for the CLAUDE.md: every fact has one home.

From Document to Harness: Guides and Sensors

An architecture document alone is not yet a harness. It becomes a harness only when we operationalize it in two forms.

Guides: The Document Steers the Agent

Guides direct the agent before the work. In the PetClinic project there are two levels:

  • The CLAUDE.md points to the architecture documents and says when to read them.
  • The skill aiup-vaadin-jooq:implement already knows the development view and applies the conventions during implementation.

The skill is essentially an executable excerpt of the architecture document. It does not just say that there is a ui and a domain package. It shows what a view looks like, what a repository looks like and how both are tested.

Sensors: The System Checks the Agent

Sensors check after the work whether the agent followed the rules. This is where the architecture document becomes something that breaks the build.

The most important sensor in the PetClinic project is ArchitectureTest, an ArchUnit test in the root package. It is the executable version of development.md. Every rule names in its because clause the document it comes from, so a failure points straight back:

@AnalyzeClasses(
        packages = "ai.unifiedprocess.petclinic",
        importOptions = ImportOption.DoNotIncludeTests.class)
class ArchitectureTest {

    @ArchTest
    static final ArchRule featuresHaveOnlyUiAndDomain =
            classes()
                    .that().resideInAPackage("ai.unifiedprocess.petclinic.(*)..")
                    .and().resideOutsideOfPackage("..core..")
                    .should().resideInAnyPackage("..ui..", "..domain..")
                    .because("development.md: each feature has exactly the sub-packages ui and domain");

    @ArchTest
    static final ArchRule noJpaNoSpringData =
            noClasses()
                    .should().dependOnClassesThat()
                    .resideInAnyPackage("jakarta.persistence..", "org.springframework.data..")
                    .because("development.md: jOOQ only, no JPA, no Spring Data");

    @ArchTest
    static final ArchRule noFetchInto =
            noClasses()
                    .should().callMethodWhere(
                            JavaCall.Predicates.target(HasName.Predicates.name("fetchInto")))
                    .because("development.md: use Records.mapping(Type::new) for compile-time column checking");

    @ArchTest
    static final ArchRule repositoriesAreNamedAndAnnotated =
            classes()
                    .that().haveSimpleNameEndingWith("Repository")
                    .should().beAnnotatedWith(Repository.class)
                    .andShould().resideInAPackage("..domain..")
                    .because("development.md: @Repository enables exception translation");

    // ... further rules for transactions, domain records, views and styling
}

The test covers the package structure, the absence of service and DTO layers, the domain-as-boundary rule, the jOOQ mapping, the persistence stereotype, the transaction boundary, session serializability and the Vaadin conventions. And the document points back to the test. The first paragraph of development.md says: most of what follows is enforced by ArchitectureTest. If you change a convention, change the rule in the same commit. A rule that contradicts the document is the rule that is wrong.

That closes the pair. The document is the readable form, the test the executable one. The agent reads both, and the build checks that they agree.

While writing the test, two things happened that show the value of this approach:

  • The rule “no cycles between features” was plausible as prose but wrong as a test. The document explicitly allows the ui of one feature to reach into the domain of another and to reference other views as routing tokens. At the feature level, the graph is therefore cyclic by design. The test only checks the domain packages for cycles. The document was imprecise, and only the test made that visible.
  • The rule “no getStyle().set()” could not be checked with callMethod(HasStyle.class, "getStyle"), because ArchUnit sees the concrete component type as the target, not the interface. The rule needs a predicate on assignableTo(HasStyle.class). A detail you only learn by running it.

A third thing happened later. The rule that a browserless test must be named *Test and a Playwright test *IT was prose in the CLAUDE.md for a while. A wrongly named test is silently never run, which is the worst kind of failure. It became TestLayerConventionsTest, two ArchUnit rules that fail the build. Whenever a rule can be expressed as a test, it should be one.

Three Sensors on the What Layer

ArchUnit checks the How layer: is the code built the way the architecture document demands? But the AI Unified Process claims more. It says docs/ is the source of truth. For a long time, that claim was not checkable. A use case could be Status: Done while an alternative flow or a business rule was never tested. A test case could be Status: Automatedwithout a journey test behind it. And when someone renamed a flow in the specification, the annotation in the test silently pointed at nothing.

Three tests close this gap, one per document family. All of them read the documents from disk, and all of them report every violation at once, so the failure reads as a work list.

UseCaseTraceabilityTest reads every docs/use_cases/UC-*.md and every @UseCase annotation on the test classpath and compares the two in both directions:

  • Referential integrity, for every use case, regardless of status. Every annotation must point at something that exists: the id at a specification file, the scenario at a ### A1: ... heading, every businessRules entry at a ### BR-NNN: ... heading. Renaming an alternative flow in the specification breaks the build until the annotations follow.
  • Coverage, only for use cases with Status: Done or Tested. Such a use case needs a test for the main success scenario, for every alternative flow and for every business rule. All other statuses are exempt, because the project writes the specification before the code, and a use case that is not yet implemented is a normal intermediate state, not a defect.
@Test
void completedUseCasesCoverEveryAlternativeFlow() {
    List<String> violations = new ArrayList<>();
    completedSpecifications().forEach(spec -> {
        Set<String> tested = testedScenarios(spec.id()).collect(toSet());
        spec.alternativeFlows().stream()
                .filter(flow -> !tested.contains(flow))
                .forEach(flow -> violations.add(spec.id() + " is '" + spec.status()
                        + "' but alternative flow \"" + flow + "\" has no test."
                        + " Annotate one @UseCase(id = \"" + spec.id()
                        + "\", scenario = \"" + flow + "\")"));
    });
    assertNoViolations("Alternative flows of completed use cases without a test", violations.stream());
}

TestCaseTraceabilityTest does the same for docs/test_cases/TC-*.md. A test case is a user journey across several use cases, in the PetClinic project for example TC-001: register an owner, find them again, review the details, add a pet, book a visit. The journey test TC001NewOwnerFirstVisitIT carries a @TestCase annotation on the class:

@TestCase(id = "TC-001", useCases = {"UC-003", "UC-004", "UC-005", "UC-007", "UC-009"})
class TC001NewOwnerFirstVisitIT extends AbstractBasePlaywrightIT {

The sensor checks that a test case with Status: Automated has a TC<NNN><Name>IT class and, the other way round, that every such class has a document, that the id agrees with the class name, and above all: that the list in useCases names exactly the use cases the Flow table of the document links, in both directions. The list is the one thing the class name cannot express, and therefore the one thing worth checking.

On top of that there is a rule across the layers: an automated test case may only walk through use cases that are themselves Done or Tested. A green end-to-end test over a use case in status Draft means one of the two status lines is lying.

BusinessRuleTraceabilityTest closes the loop between docs/business_rules.md and the use cases. Every GR-NNN a use case references must exist. The Realized by: line of a shared rule must name exactly the use case rules that reference it. The summary table must match the rules below it. And a rule realized by only one use case does not belong in the catalogue at all. Nothing in Markdown enforces any of this. The test does.

The consequence is stated in testing.md: the Status: line is an assertion, not a label. Setting it to Done or Automatedswitches the sensor on. That is exactly why it must not be set without a coverage audit first. Hooks now enforce that; more on that below.

Three details of the tests are worth noting:

  • Every annotation has one place. @UseCase sits on the method, because a use case has many coverage units (main scenario, each flow, each business rule) and only the author knows which method covers which unit. @TestCase sits on the class, because a test case has exactly one unit: the journey. Both sensors reject their annotation anywhere else. Without this rule, a journey test could silently “cover” one of the five use cases it walks through, which would look like coverage while the alternative flows of that use case stay untested.
  • A guard test checks that specifications and annotations were found at all. Without it, a wrong working directory would let every check pass over an empty set, a sensor that reports green because it is blind.
  • The tests report every violation at once, not just the first. The failure report is therefore a work list, and one the agent can work through directly.

The other sensors in the project:

  • jOOQ generates code from the schema. A wrong column name breaks compilation. The Flyway migration is therefore the schema DSL, and the entity model must match it.
  • Browserless tests (UC<NNN><Name>Test) check every view against the use case specification. Every test method carries a @UseCase annotation with id, scenario and business rules. That is the input for UseCaseTraceabilityTest.
  • Playwright tests (TC<NNN><Name>IT) check whole user journeys in the browser. Every class carries a @TestCaseannotation. That is the input for TestCaseTraceabilityTest.
  • TestLayerConventionsTest makes sure the naming convention that decides the build phase (*Test runs in test, *IT in verify) is followed, so no test is silently skipped.

The difference between guide and sensor is the difference between probabilistic and deterministic compliance. A guide raises the probability that the agent works correctly. A sensor makes sure that incorrect work does not get through. A good harness needs both.

Hooks: Sensors for the Session

All the sensors above share a limitation that is easy to miss: they are post-hoc. ArchUnit and the traceability tests can only catch drift once somebody runs them. Nothing in the source tree can express “the agent ran the sensors before it said it was done”, because that is a property of the session, not of the repository.

So one failure mode survives every test in the project. The agent edits a view, skips ./mvnw test, and reports success. The drift is real from that moment until CI catches it, minutes or hours later, and only if someone pushes.

A second gap has the same shape. The CLAUDE.md says that a Status: line is an assertion and must not be set without a coverage audit. Prose shapes behavior probabilistically. It does not stop an edit. The sensor does catch a wrong status, but only on the next full run, and a narrowed run (-Dtest=...) skips the sensors entirely while still reporting green.

Claude Code hooks close both gaps. They are shell scripts that Claude Code runs on events of the session: when a session starts, before or after a tool call, when a subagent finishes, before a turn ends. A hook that exits with code 2 blocks the action and hands its message back to Claude.

The First Version, and What a Review Found

The first version had four hooks. One recorded the commit the session started on. One watched Edit and Write calls and flagged a Status: line that changed value. One watched Bash calls and, when a full ./mvnw test ran and the console showed no failure, remembered that the sensors had run. And the Stop hook refused to end a turn while src/ or docs/had changed since that run.

A review showed that each of them could be satisfied while the property did not hold, and always for the same reason: the hook keyed on which tool ran or on what the console printed, not on the state of the repository.

  • The status guard matched Edit and Write. A sed or a heredoc through Bash never reached it. And in auto mode Claude Code prefers Bash for edits, so the bypass was the default path, not an edge case.
  • The sensor recorder took any full run whose output lacked failure text as green. A run in the background, piped through tail, redirected to /dev/null or followed by || true produced no failure text and wrote the marker without a passing build.
  • The status guard fired after the edit. The claim was already in the file, and “run the audit first” was a request, not a rule.
  • The only run that counted was the full Testcontainers suite. Every turn that touched a document cost a container start, and “Docker is not available” was a sanctioned way to end a turn unverified.

That is the same lesson as with the ArchUnit rules, one level up. A rule that is plausible as prose is often wrong as a check, and you only find out by trying to break it.

The Second Version: State and Evidence

The second version follows two rules, written down in ADR-011.

A guard reads the state of the repository, never the shape of a tool call. check-spec-status.sh runs after every tool call, rereads every Status: line in docs/, and compares it with what it saw last. Which tool made the change is irrelevant. A PreToolUse guard for Edit and Write stays, because there it can refuse the edit before it lands. It is the fast path, not the guarantee.

Evidence is positive. A sensor run counts because the Surefire report of every sensor class exists in target/surefire-reports/, is newer than the session start and than every changed file, ran tests, and counts no failure or error. The test JVM writes those files whatever the console shows. An audit counts because the uc-coverage agent of the AIUP plugin finished, which Claude Code reports as a SubagentStop event. The guards accept Done, Tested or Automated only with such an audit marker newer than the last change under src/. The marker says an audit happened, not that it found nothing: the sensors judge the claim itself on the next run.

This is what the six hooks look like now:

HookEventWhat it does
session-start.shSessionStartrecords the commit the session started on and when, drops the previous session’s markers, and takes a first reading of every Status: line
guard-spec-status.shPreToolUse(Edit, Write)refuses an edit that sets a Status: to Done, Tested or Automated unless a coverage audit of that specification finished in this session after the last change under src/
check-spec-status.shPostToolUse(every tool)rereads every Status: line after each tool call and reports one that now asserts coverage without an audit behind it, whichever tool changed it
record-coverage-check.shSubagentStop(uc-coverage)records that the coverage agent audited a UC or TC: the evidence the two guards above ask for
record-sensor-run.shPostToolUse(Bash)after a Maven test or verify, records the change set the sensors saw, if and only if every sensor’s Surefire report is fresh and green
require-sensors.shStoprefuses to end a turn while src/, docs/ or pom.xml changed and the sensors have not run since the newest change, or while a hook changed and smoke.sh has not

The Stop hook is the one that justifies the layer. When Claude wants to finish a turn, the hook compares the changed files against the change set the last green run saw. If a file was edited after that run, or added, deleted or committed and the run did not see it, the hook returns exit code 2 with a message that says what to run. Claude reads the message, runs the sensors, and tries to stop again.

Making the Sensors Cheap Enough to Run Every Turn

The Docker excuse was closed by removing the reason for it. The five sensor classes (ArchitectureTest, TestLayerConventionsTest and the three traceability tests) carry the JUnit tag sensor, and a Maven profile activated by -Dgroups=sensor skips jOOQ code generation and the JaCoCo agent for that run:

./mvnw -q test -Dgroups=sensor

Seconds instead of minutes, no container. That is the run the Stop hook asks for. The full suite remains the definition of done and the commit gate; a full test or verify counts as well.

Three more decisions kept the layer usable:

  • A hook must fail open. Unparseable input, a missing .git, no jq: the hook exits quietly. A guardrail that breaks the session is worse than the drift it was meant to catch.
  • The Stop hook is one enforced reminder per turn, not a wall. When Claude tries to stop the second time, Claude Code sets stop_hook_active, and the hook lets the turn end. A turn that ends without the sensors having run is visible in the transcript as exactly that. This is the one deliberate escape. The other one from the first version, touching the marker file by hand, no longer works, because the hook reads the Surefire reports behind the marker.
  • A commit must not clear the guard. The Stop hook diffs the working tree against HEAD, which a commit would silently reset. That is why session-start.sh records the commit the session began on, so the hook can also diff HEAD against that commit.

The hooks themselves are watched too. smoke.sh beside them exercises all six against a throwaway repository and records its own green run. A change under .claude/hooks/ or to .claude/settings.json holds the turn until smoke.shhas run, and CI runs it before the Maven build. It is a shell script and not a Maven test on purpose: the hooks are a property of a Claude Code session, and this is the closest thing to one.

Finally, .claude/settings.json allows the Maven build, the smoke test and read-only git commands, so a fresh clone can satisfy the hooks without a permission prompt. Otherwise the hook tells Claude to run a build that Claude then has to ask permission for.

The Dividing Line Between Hooks and Tests

The hooks were tempting to extend. Why not check the test naming rule in a hook too? Because that rule can be checked by reading the repository, and so it became TestLayerConventionsTest instead. The rule for adding a hook is written down in ADR-010 and ADR-011: if a rule can be checked by reading the repository, it is a test. Only a rule about what happened during a session is a hook.

The reason is portability. Hooks live in .claude/, fire only inside a Claude Code session, and run for nobody else: not CI, not another editor, not a human contributor. An ArchUnit rule binds everyone and travels with the clone. Hooks buy latency, tests buy correctness. Only CI decides whether code is allowed to exist.

So the harness in the PetClinic project has four enforcement layers: the CLAUDE.md as the map, the architecture views and the skills as guides, the ArchUnit and traceability tests as sensors on the repository, and the hooks as sensors on the session. The last layer is the smallest, and it should stay that way.

What This Means in Practice

The CLAUDE.md is a map, not a library. It contains what the agent cannot find anywhere else: the source of truth, the stack, the commands, and when to read which document. Everything else belongs in separate documents that are read on demand. The ETH study shows that more context does not bring more quality, only more cost. And the CLAUDE.md grows on its own if you do not trim it regularly.

The Software Architecture Document is a first-class artifact again. Split into the 4+1 views, each short and concrete, with code examples, and each answering one question. Together they answer the questions the agent would otherwise fill with generic answers from training.

Sensors belong on both layers. ArchUnit checks that the code fits the architecture. The three traceability tests check that the tests fit the specification and that the specification agrees with itself, for use cases, test cases and business rules. Only with both is the claim “docs is the source of truth” more than a sentence in the CLAUDE.md.

Architecture decisions should be implemented as sensors. Every rule that can be checked with the compiler, ArchUnit or a test should be checked. And the document should point to the test, and the test to the document. Everything else remains a hope.

Tests guard the repository, hooks guard the session. A test checks the source tree. A hook checks what happened in a session: that the sensors actually ran, that a status was audited before it was set. A hook reads the state of the repository and demands positive evidence, never the shape of a tool call. Keep the hook layer small; whatever can be a test should be a test.

Harness Engineering is architecture work. The role of the architect shifts from code reviewer to designer of the system in which the agent works. Whoever describes and enforces the architecture clearly gets code that fits the system.

The AI Unified Process makes this connection explicit. The What layer says what is built. The architecture documents in the Harness layer say how it is built. The sensors make sure it was built that way. And the hooks make sure the sensors were asked.