Skip to content

Reference

Who is this for? Someone asking what a check catches, what it prints, and why it exists at all. When should I read it? When choosing a check, or when one just fired. To turn one on, the guide's table; for every config key, configuration.


One section per situation, with the output the tool really prints — none of it is a mock-up. Each names the check to turn on and where the rule came from; the reasoning behind the model as a whole is the standard.

Check It fails when
proof a documented claim's file, test or command is gone
tracks a work track links a handoff that was deleted
tasks a ticked checkbox names no check
queue a task claiming to pass fails when re-run
locked an agent commit touched the files that grade it
cold-start a fresh clone cannot install and verify itself
clean-exit a session left debris, or wrote nothing down
instructions the instruction file grew past its line limit
release the manifest, the changelog and the tags disagree

1. The README describes software that does not exist

Comes from: Lecture 03 — the repository is the source of truth. Not from the course: the marker syntax itself, which came out of production repositories.

Brownfield, inherited repo, or an agent that documented its intentions.

You write a claim once and attach its evidence:

Totals round per currency, not per line.   <!-- proof: docs/<a-file>.md -->
Check it yourself with `make verify`.     <!-- proof: make <a-target> -->

(Real markers name a real path; the angle brackets above are only so this page's own examples are not checked as claims.) Delete the document, or rename the target, and the build says so:

FAIL  proof markers
  README.md:31  docs/ARCHITECTURE.md
      no such file
  README.md:32  make verify
      no make command named "verify"

Turn on: docs. Add the documents that may never drift to docs.mustCarryProof — those must carry at least one marker, so a rewrite cannot quietly drop the evidence.

2. "Works on my machine" — and only there

Comes from: Lectures 03 and 10 — the repository is the source of truth, and the real run is the proof.

Onboarding, a new agent session, a fresh CI runner.

cold-start clones your repository into an empty directory and runs the commands your own documentation gives a newcomer. No caches, no environment carried over, nothing that only exists on the machine that built it.

FAIL  cold start
  the clone ran: npm ci && npm test
    npm ci failed: package-lock.json is not committed

Turn on: coldStart, listing requiredFiles, entryDocs and the commands a newcomer runs.

11. A quality number that quietly slipped

Comes from: Not from the course. Both repositories this package was extracted from invented this rule independently and wired it by hand.

The evals still run. The number is a little lower every month, and no single month is the one anybody notices.

One repository keeps evals/thresholds.json — recall, groundedness, citation validity, each with a floor — and fails CI when a score drops below one. The other has a grounding gate with a pass mark. Both put the floors file out of the agent's reach, and that is the load-bearing half: a loop that can lower its own pass mark grades itself.

FAIL  thresholds
  evals/results.json:1  recall_at_5
      0.71 is below the floor of 0.8 declared in evals/thresholds.json
    fix:  raise the score. Lowering the floor to pass is the move this check exists to catch

It does not run your scoring. Your command writes the results file; this reads it and compares. A metric with a floor that nobody measured fails too — a scoring step that did not run is not a scoring step that passed.

Turn on: thresholds, naming results, floors and the command that produces the results. If the repository already declares locked.paths and the floors file is not among them, the check says so. Floors only: a metric where lower is better is the same rule mirrored, and inventing a direction on zero real cases would be a guess.


3. A feature that takes four sessions

Comes from: Lecture 05 — continuity between sessions. The handoff shape is not from the course.

Context dies at the end of every session; the next one re-derives it, or re-decides a question that was already settled.

Live work lives in specs/TRACKS.md — one line per track, each pointing at a handoff.md that says what to load, what not to load, where the work stopped and what the first step is. The next session gets it handed over at startup rather than re-reading the repository:

$ harnessimo brief
== specs/TRACKS.md (live work tracks — load a track's handoff first) ==
- Checkout totals — [handoff](specs/0007-checkout/handoff.md) — active, next: T3 currency rounding
- Hosted deploy — none — blocked (on an API token)

The index is machine-checked, because a handoff that lies is worse than no handoff: a track with no status, or a link to a handoff someone deleted, fails the gate.

Turn on: tracks.

4. "Done" that stopped being true

Comes from: Lectures 08 and 09 — feature lists as primitives, and declaring victory too early.

The task was genuinely finished in March. Something unrelated broke it in May, and the board still says done.

Each queue item carries its own verification command. Only the tool writes passing, and only after that command exits 0 — and --reverify re-runs every one of them in CI:

FAIL  queue
  "checkout-totals" claims passing but its verification fails now
    command: npm test -- checkout
    fix:  fix the regression, or the claim was never true

Editing state by hand is not a shortcut; it is the thing this check catches.

Turn on: queue.

5. An agent that edits its own exam

Comes from: Not from the course. Locked surfaces came out of production, where an agent edited the script that graded it.

The fastest way to make a red build green is to change what "green" means.

Name the surfaces that define success — the CI workflows, the constraints file, the eval thresholds. A commit carrying the agent trailer that touches them fails the build:

FAIL  locked surfaces
  a1b2c3d  Co-Authored-By: Claude  touched .github/workflows/ci.yml
    fix:  a human makes this change, or the baseline records why it moved

This is drift detection in CI, not a sandbox — it catches the honest case, which is the common one.

Turn on: locked, with paths, a baseline and your agentTrailer.

6. The long unattended run

Comes from: Lecture 12 — a clean state at session end.

You start it and go to bed.

clean-exit reads what the session actually changed and refuses the two ways a run ends badly without failing: debris left in the code (TODO, debugger, .only( — you choose the markers), and a progress file that stayed frozen while hundreds of lines changed around it.

FAIL  clean exit
  src/checkout/totals.ts:88        debugger
  .harness/4-state/PROGRESS.md     unchanged while 412 lines of code moved
    fix:  write down where this got to, or the next session starts blind

Turn on: cleanExit. The progressThreshold keeps a one-line fix from tripping it.

7. "What does this repo actually enforce?"

Comes from: Lecture 04 — one giant instruction file fails; the honesty of doctor is this project's own rule.

A new contributor, a code review, or you six months later.

$ harnessimo doctor
  enforced      proof markers    docs claims resolve to real files, tests and commands
  enforced      work tracks      the track index resolves and every track carries a status
  not enforced  queue            no queue section in harnessimo.config.json

A check you did not configure is reported as not enforced, in the same output, with the same weight. A harness that overstates its own coverage is the failure it exists to prevent.

8. The version means different things in different places

Comes from: Not from the course, and not from anywhere else: this repository published three versions past its own last tag.

The badge says one thing, the registry serves another, and both are right about themselves.

This one is not hypothetical: while writing this tool, three versions reached npm through a manual workflow run while the repository's own tags stopped two releases earlier. Nothing noticed, because nothing was checking the most duplicated claim a project makes.

FAIL  release
  CHANGELOG.md:1  v0.4.3
      0.4.3 is described as released and has no v0.4.3 tag
    why:  a version published outside the release flow leaves the repository behind
    fix:  cut the release, or remove the entry if it never shipped

Turn on: release, naming your manifest, your changelog and the tag prefix. It reads the version from the manifest, the entries from the changelog and the tags from git; a version nobody wrote down, an entry that is not at the top, or a released version with no tag all fail.


9. The session is expensive and nobody is counting

Long agent runs, where most of the budget goes before anything is finished.

The largest avoidable cost in a long session is the same file read twice. The second read costs its full length again and teaches the model nothing it does not already hold — and nothing reports it, so nobody fixes it.

$ harnessimo budget
harnessimo budget — this session (estimated at 4 bytes per token)

  read          38 file(s), ~184k tokens
  re-read        9 file(s), ~41k tokens — 22% of everything read
  largest     src/pipeline.ts, read 3×, ~7k each

  fix:  a re-read is a session that lost its place. `harnessimo guard read <path>`
        refuses the second read of a file that has not changed.

Wired as a tool-use hook, the guard answers before the read happens: allowed the first time, refused the second if nothing changed, allowed again the moment it does. A range is not a whole file, and --override always wins — and is counted, so a rule that gets in the way shows up in the ledger rather than in somebody's frustration.

Turn on: tokens. Note that this is not a tenth check: the nine answer "is this finished", and this one fires while the work happens. doctor lists it apart for that reason, and check does not run it — a completion gate that depended on session state would be a gate nobody could reproduce.


10. A key, a colour or an import that belongs somewhere else

Comes from: Not from the course. Both repositories this package was extracted from wrote this rule by hand, differently, because there was nowhere to declare it.

The rule is in someone's head, or in a script, or in a review comment nobody made this time.

One repository forbids the service-role key anywhere under app/ — it bypasses row-level security, so a client-side use is a data leak wearing an ordinary import. It also forbids hard-coded colours in components, because design tokens live in one file. The other forbids placeholder markers in published content. Three rules, one shape: this pattern, not in these files, for this reason.

FAIL  boundaries
  app/dashboard/page.tsx:14  service_role
      the service-role key bypasses row-level security; it belongs on the server

The reason is required, and it is the whole point. A boundary that prints "matched /service_role/" says what happened; the next person needs to know what to do instead.

Turn on: boundaries, listing rules of pattern, paths and reason (plus allow for the one legitimate place). This repository runs one on itself: an import in src/ that is neither a node: builtin nor relative is a runtime dependency, which its constraints forbid — a rule that was prose a reviewer had to remember, and is now a line that fails.


A full session, end to end

$ harnessimo brief
== specs/TRACKS.md (live work tracks — load a track's handoff first) ==
- Checkout totals — [handoff](specs/0007-checkout/handoff.md) — active, next: T3 currency rounding

== work queue ==
active: checkout-totals — an order total matches the sum of its lines, in every currency
  verify with: harnessimo queue verify checkout-totals

== enforced here ==
proof, tracks, tasks, queue, coldStart, cleanExit

# ... the agent works ...

$ harnessimo queue verify checkout-totals
running: npm test -- checkout
  ✓ 12 tests passed
checkout-totals: passing — evidence recorded

$ git commit -m "feat(checkout): per-currency rounding (0007)"
harnessimo: proof markers ok · tracks ok · task gate ok · queue ok
[main 4f1a2b9] feat(checkout): per-currency rounding (0007)

The commit went through because the checks passed, not because anybody said it was done.


Next: what was taken from OpenSpec, Spec Kit, BMAD and the rest, or adopting it in an existing repo.