“All tests pass. The use case is done.” Every developer who works with an AI agent has read this sentence. And every one of them has found out at least once that the tests never ran. The agent edited a view, skipped the build and reported success. Not out of malice. Saying “done” is cheaper than checking it.

You can write in the CLAUDE.md that the agent must run the tests before it finishes. Most of the time it will. But a rule in the CLAUDE.md is a wish. A failing build is a fact. This post shows how Claude Code hooks turn “done” from a claim into something the session has to prove, with examples from the PetClinic of the AI Unified Process.

“Done” Is Two Claims

When the agent says it is done, it claims two different things:

  • The checks ran, and they were green. This is a property of the session. Nothing in the repository can say whether the agent ran the tests before it ended its turn.
  • The specification is fulfilled. In the PetClinic, each use case has a Status: line. Done or Tested is not a label. It switches on a sensor, UseCaseTraceabilityTest, that from then on demands a test for the main success scenario, every alternative flow and every business rule.

Tests can check the second claim, but only when somebody runs them. Nothing can check the first claim except the session itself. That is the job of a hook.

What a Hook Is

A Claude Code hook is a command that Claude Code runs on an event of the session: when the session starts, before or after a tool call, when a subagent finishes, and when Claude wants to end its turn. The hook gets the event as JSON on stdin. If it exits with code 2, Claude Code blocks the action and gives the hook’s stderr to Claude as feedback. Claude reads it and carries on.

Hooks are configured in .claude/settings.json and committed with the code, so every clone gets the same guardrails. This is the hook section of the PetClinic, shortened:

{
  "hooks": {
    "SessionStart": [
      { "matcher": "startup|clear",
        "hooks": [{ "type": "command", "command": "\"$CLAUDE_PROJECT_DIR/.claude/hooks/session-start.sh\"" }] }
    ],
    "PreToolUse": [
      { "matcher": "Edit|Write",
        "hooks": [{ "type": "command", "command": "\"$CLAUDE_PROJECT_DIR/.claude/hooks/guard-spec-status.sh\"" }] }
    ],
    "PostToolUse": [
      { "matcher": "Bash",
        "hooks": [{ "type": "command", "command": "\"$CLAUDE_PROJECT_DIR/.claude/hooks/record-sensor-run.sh\"" }] },
      { "hooks": [{ "type": "command", "command": "\"$CLAUDE_PROJECT_DIR/.claude/hooks/check-spec-status.sh\"" }] }
    ],
    "SubagentStop": [
      { "matcher": "aiup-vaadin-jooq:uc-coverage$",
        "hooks": [{ "type": "command", "command": "\"$CLAUDE_PROJECT_DIR/.claude/hooks/record-coverage-check.sh\"" }] }
    ],
    "Stop": [
      { "hooks": [{ "type": "command", "command": "\"$CLAUDE_PROJECT_DIR/.claude/hooks/require-sensors.sh\"" }] }
    ]
  }
}

Six hooks, but they serve two ideas: the Stop hook and the status guard. Everything else collects the evidence those two need.

The Stop Hook: No End of Turn Without a Green Run

The Stop hook runs when Claude wants to end its turn. In the PetClinic, it refuses while src/, docs/ or pom.xml changed after the last green run of the sensors. Claude gets a message that says what to run, runs it, and tries to stop again.

The real script handles commits during the session, deleted files and worktrees. Here is a simplified version with the same core that you can drop into your own project. Adjust the report names to your test classes:

#!/usr/bin/env bash
# Stop: no end of turn while code or specs changed after the last green test run.

# Fail open: without jq this hook cannot tell a first stop from a second one.
command -v jq >/dev/null 2>&1 || exit 0
input=$(cat)

# The second stop after a block goes through, so a session that cannot build
# does not loop forever.
[ "$(jq -r '.stop_hook_active // false' <<<"$input" 2>/dev/null)" = "true" ] && exit 0

cd "${CLAUDE_PROJECT_DIR:-.}" 2>/dev/null || exit 0
changed=$(git status --porcelain -- src docs pom.xml 2>/dev/null | sed 's/^...//') || exit 0
[ -n "$changed" ] || exit 0

problem=""
for name in ArchitectureTest UseCaseTraceabilityTest; do
    report="target/surefire-reports/TEST-com.example.$name.xml"
    if [ ! -f "$report" ]; then problem="there is no test report for $name"; break; fi
    if grep -q 'tests="0"' "$report"; then problem="$name ran no test"; break; fi
    if grep -qE '(failures|errors)="[1-9]' "$report"; then problem="$name failed"; break; fi
    while IFS= read -r file; do
        if [ "$file" -nt "$report" ]; then problem="$file changed after the last run"; break 2; fi
    done <<<"$changed"
done
[ -z "$problem" ] && exit 0

cat >&2 <<MSG
You are not done: $problem.
Run ./mvnw -q test in the foreground and let it finish. Console output
does not count, only a fresh and green test report.
MSG
exit 2

Register it under "Stop" in .claude/settings.json, and the agent can no longer end a turn on a changed tree without a fresh, green report.

Note the stop_hook_active check at the top. When Claude tries to stop a second time after a block, Claude Code sets this flag, and the hook lets go. Without it, a session that cannot build, for example because a database is down, would loop forever. So the Stop hook is one enforced reminder per turn, not a wall. A turn that still ends without a green run is visible in the transcript as exactly that.

Positive Evidence: The Report, Not the Console

The obvious way to know whether the tests ran is to look at the output of the last ./mvnw test. The first version of the PetClinic hooks did exactly that, and a review showed how easily it is fooled. A run in the background, a run piped through tail, redirected to /dev/null or followed by || true: none of them prints failure text, and all of them looked green. A narrowed run with -Dtest=… really is green, but it skips the sensors.

So the hooks ask for positive evidence. A run counts only if the Surefire report of every sensor class:

  • exists in target/surefire-reports/,
  • is newer than the session start and than every changed file,
  • counts at least one test,
  • and counts no failure and no error.

The test JVM writes these files itself, whatever the console shows. A skipped run writes none. A failed run says so in the file. A run that is still going in the background has not written them yet. A narrowed run leaves the reports of the other sensors stale. And touch on a marker file proves nothing, because the hook reads the reports behind the marker.

Evidence must also be cheap, or the agent will look for reasons to skip it. The five PetClinic sensors carry the JUnit tag sensor, and a Maven profile skips code generation and coverage for them. ./mvnw -q test -Dgroups=sensor takes seconds and needs no Docker. That removed the most common excuse: “Docker is not available, so I could not run the tests.”

The Status Guard: Done Only After an Audit

The second claim is the Status: line. Setting a use case to Done by hand either breaks the build, or, worse, certifies coverage that was never checked. In the PetClinic, Done, Tested or Automated is accepted only after the uc-coverage agent of the aiup-vaadin-jooq plugin has audited that specification in this session, after the code and tests last changed. You start the audit with /coverage-check UC-003. It maps every step, flow and business rule of the use case onto the code and the tests and reports the gaps.

Three hooks work together for this:

Hook Event What it does
record-coverage-check.sh SubagentStop When the uc-coverage agent finishes, writes a marker for each audited id. This is the evidence.
guard-spec-status.sh PreToolUse (Edit, Write) Refuses an edit that sets an asserting status without a fresh audit marker, before the edit lands.
check-spec-status.sh PostToolUse (every tool) Rereads every Status: line after each tool call and reports one that changed without an audit, whichever tool changed it.

Why two guards? The PreToolUse guard can stop the edit before it happens, but it only sees Edit and Write. A sed or a heredoc through Bash goes past it, and in auto mode Claude Code often prefers Bash for small edits. So the second guard ignores the tool call completely and reads the files. A guard that reads the state of the repository holds whichever tool made the change.

This is what Claude sees when it tries to skip the audit:

Refused: UC-003-register-new-owner.md would read "**Status:** Done".

That line is an assertion the traceability sensors act on, not a label, and no
coverage audit of UC-003 has finished in this session since the code and tests
last changed. Run

  aiup-vaadin-jooq:coverage-check UC-003

first. When it reports no gaps, make this edit again and then run the sensors
(./mvnw -q test -Dgroups=sensor) so the one it switches on actually votes. When
it reports gaps, close them or leave the status as it is.

The message does not only say no. It says what to do next, so the agent can fix the problem in the same turn. And the marker only says that an audit happened, not that it found nothing. Whether the claim holds is decided by the sensors on the next run.

Fail Open

A hook runs inside every session. If it breaks, the session breaks. So every hook in the PetClinic fails open: unparseable input, a missing .git directory, no jq on the machine, and the hook exits 0 and says nothing. You can see it in the sketch above: command -v jq || exit 0 and || exit 0 after every command that might fail.

This sounds like a hole, and it is one. But a guardrail that blocks a session because of its own bug does more damage than the drift it was meant to catch. Developers disable hooks that get in their way, and then nothing is checked. CI is still there as the last line.

Fail open does not mean untested. The PetClinic has a smoke.sh next to the hooks that runs all of them against a throwaway git repository: a change without a run blocks, a failed report blocks, a stale report blocks, a commit does not clear the guard, a sed to Done is reported. And a change to a hook holds the turn until smoke.sh is green, so the guards guard themselves too.

Hook or Test?

Once hooks work, it is tempting to put every rule into one. Don’t. Hooks run only inside a Claude Code session. They do not run in CI, in another editor or for a developer without Claude Code. A test binds everybody and travels with the clone.

So the rule of thumb is: if you can check it by reading the repository, write a test. Only rules about the session become hooks. “A browserless test must be called *Test” is a property of the repository, so it became an ArchUnit test. “The agent ran the sensors before it said it was done” is a property of the session, so it is a hook. Hooks buy latency, tests buy correctness. Keep the hook layer small.

Try It Yourself

This is lab 11 of my Spec-Driven Development workshop:

  1. Clone the PetClinic. It already has specifications, sensors and hooks.
  2. Run ./mvnw test -Dgroups=sensor and look at the reports in target/surefire-reports/.
  3. Rename an alternative flow in one of the docs/use_cases/UC-###.md files and run the sensors again. UseCaseTraceabilityTest tells you which annotation now points at nothing.
  4. Start Claude Code and ask it to set a use case to Done. Watch the hooks refuse it, and watch Claude run /coverage-check to earn the status.
  5. Read ADR-007, ADR-010 and ADR-011 in docs/architecture/adr/. They explain why each sensor and each hook exists.

If you want the full story of how these hooks came to be, including the first version and what the review found, read Harness Engineering: Why a Minimal CLAUDE.md and a Good Architecture Document Belong Together.

Conclusion

An agent that says “done” is making a claim. Hooks let the session ask for proof: the Stop hook wants a fresh, green test report, and the status guard wants an audit before a use case is marked as done. Both read the state of the repository, not what the agent says or what the console prints. Both fail open. And both stay small, because whatever can be a test should be a test.