Skip to content
3 min read overview, plus optional history

Command

This page describes Command as of September 11, 2026. The kit was one day into its trial.

What it is

Command is a shared kit for my Claude Code project sessions: skills for checking work and reaching me, rules about what a session may do, and hooks that check actions such as merges. The idea is to carry those tools into the project where the work happens, so I can work directly with a session there. Command’s private repository holds the kit’s source and the goals file those sessions work toward [measured 2026-09-11, the kit plan and the goals file].

What a session inherits

PartUse in a project session
Instructions and output styleSet the working agreements and how the session reports back
SkillsCheck status, bring a decision to me, or run a retrospective when asked
Repository listFind the project, its merge restrictions, and the label that reaches me
Guard hooks and settingsCheck actions against the rules; the installed merge hook required a reviewed commit

For example, a session doing engineering work can use the status skill to check the repository, the review loop to examine its changes, and the ask skill when it needs my decision. These are shared tools available from the project, rather than work handed out by a separate dispatcher [measured 2026-09-11, the kit plan’s contents and keep list].

The trial

The first phase runs from September 10 to 17, with the dispatcher off and nothing deleted. It installs the status, ask, and retro skills. The plan tests whether project sessions can carry the work with acceptable token use, no ownerless work, fewer status questions from me, and a weekly venture release. Retiring the old system is a second decision, after the trial [measured 2026-09-11, the kit plan and the goals file].

Why I’m trying the simpler design, the strongest objection, and the trial’s measures.

Earlier operator reference

Before the kit, Command coordinated work through a manager layer and then a board with a dispatcher. The reference below describes those earlier designs. Its counts and operating claims belong to the dates shown, not to the kit today.

Read the August 16 operator reference and updates through September 11, 2026

Since August 24

The reference below is the office shape. It closed on September 3, and on September 10 the kit shape began a seven-day trial. Each row says what moved and what carried forward, as read on September 11.

Bounded update
Merge control. The auto-merge workflow that the green tier rests on was disabled by hand on August 31 and is still off; a session found it on September 2 with seventeen pull requests open [measured 2026-09-11, the repository’s workflow list and its last run]
The office. The manager layer closed on September 3. Its six scheduled loops, the morning letter, the nightly shift, the four-hourly remediation tick, the desk heartbeat, the wiki digest and the wiki lint, are all disabled [measured 2026-09-11, the scheduler]. A board of files and a dispatcher with no model in it replaced it
The dispatcher. Turned off on September 9 after it launched a worker against finished work; its Windows task is disabled [measured 2026-09-11, the scheduled-task state]
The always-on service. Stopped with the office [measured 2026-09-11, no process on its ports]. The 69-repository view in the August reference is the registry’s entry count, last changed August 6, not a running poll
The three channels. The daily letter is paused with the office. A decision now reaches me as a GitHub label and a desktop notification, and the count of asks that expire unanswered is not yet measured [measured 2026-09-11, the kit plan]
The kit plan. Approved September 10. Phase A installs the first three kit skills and runs to September 17 with nothing deleted; phase B, on a second go from me, retires the board, dispatcher, desk file, backend and wiki and archives the charter [measured 2026-09-11, the kit plan]
What carried forward. The goals file, the weekly worktree sweep, the guard hooks that refuse a merge past a hold, the authority colors, and the loop-card discipline below, which the kit carries as hooks and skills rather than as an office [measured 2026-09-11, the kit plan’s keep list]

Since August 16

The broader inventory on August 21 found 45 registered automatic mechanisms [measured 2026-08-21, the mechanism registry] and 29 scheduled jobs, 18 enabled and 14 recurring [measured 2026-08-21, the scheduler inventory]. Six of those recurring jobs belong to Command [measured 2026-08-21, the recurring-job ownership audit]. This is a broader inventory than the Command-only scheduled-jobs and loop-registry tables in the August 16 reference, so the newer counts are not directly comparable to the older ones.

Bounded update
Jobs. Three of Command’s six recurring jobs produced durable output, one could not be measured, one was open at its acting end by design, and one was failing [measured 2026-08-21, the recurring Command job audit]
PRs. Two remediation pull requests that had finished but stalled behind unanswered review threads both merged on August 23 [verified 2026-08-23, the remediation pull-request records]
Ops. The parked checkout returned to main, two decision-board fact-checkers resumed, and an automatic re-entry brief was added [verified 2026-08-24, the operating records and decision-board probes]
Plan. The repair that gives each report-only output a named consumer, or retires it with a date, is specified but not built [verified 2026-08-24, the named-consumer repair specification]

The reference below describes Command as it stood on August 16, 2026. Most of it was younger than that by weeks: the charter and the authority rules were ratified on July 22, and the registry of unattended jobs landed on August 5. Every count in that reference was re-derived on August 16 against the live repository, and says where it came from.

The always-on service

One thing in Command is long-running, and it only senses. A local service polls all 69 repositories on a three-minute cycle, reading git state, CI results, service probes and the trail my sessions leave behind, and keeping a per-project metric history about a day deep. No model runs in that loop. It ranks what needs attention, it serves the live map that is my inspection surface, and it never acts.

Nothing else here is a long-running process, and nothing long-running decides anything. Deciding happens in sessions, and in a set of files they all read.

Roles

Three of these are the cabinet: defined staff roles invoked where an independent read matters. They are prompts with written criteria, not personas talking to each other.

RoleWho holds it, and what it does
OwnerMe. Sets direction, signs the handful of things only an owner can sign, and can open any part of it at any time. Owning is not watching: nothing in the design depends on me remembering anything
PresidentA role, not a process, held by whichever work session is currently on duty. Holds the company’s live state between sessions, in a file rather than in that session’s head, so a session that starts tomorrow is the same President as the one that stopped tonight
AllocatorA cabinet role. Ranks the work
ControllerA cabinet role. Keeps the ledgers honest
AuditorA cabinet role. Reviews adversarially, and holds the second key on anything being reclassified
Portfolio companiesThe projects themselves, each with its own sessions and its own context. They do the work. Command directs them by filing a requirements issue that states the what and the why, and never opens their files to write the how

The desk file

The President’s memory is the desk file: a short, bounded page of live threads, open judgments, watch-items, and recent decisions with a line of reasoning each. A session reads it when it clocks in and writes it when it clocks out.

One line in it names the single session currently holding the role. That claim exists because announcing a handoff doesn’t perform one: spawning a successor session does not stop its predecessor, so “handed off” and “still running” were the same state. On August 12 two sessions were live at once and my decisions landed in the one that had already handed off. The successor found the claim, stood down, and lost nothing, but two sessions acting on one decision is what it risked.

The three channels

Everything between the company and me flows through exactly three channels, and the count is the point.

ChannelWhat it carriesWhich way it moves
The daily letterOne letter a dayPushed at me
The decision folderDecisions, each shaped so it is answerable in a wordPushed as they arise
InspectionAnything, at any depthI pull it

Whatever I find on the way down becomes a work item for staff, never an obligation I carry back out. Curiosity is not allowed to convert into homework.

Outcomes, standing rules, and observations

Three kinds of durable state sit under all of it.

KindWhat it is
OutcomesWhat I fund with attention. 70 live [measured 2026-08-16, the objective store], each carrying an autonomy dial and a state that survives across sessions
Standing rulesRules that reach the whole portfolio, and only by walking candidate to piloting to proven, so nothing lands on 69 projects on a hunch
Ranked observationsRecurring friction, ranked

The third of those is the input to the self-improvement loop, which the “what doesn’t work” section covers, because it hasn’t closed yet.

Authority colors: green, yellow, red

Every piece of work gets a color when it enters the queue, from a written checklist and not from a judgment call. The rulebook that defines them is called the Delegation of Authority.

ColorWhat staff may doWho releases the change
GreenJust does itStaff, without asking me
YellowBuilds it completely: tests and pull request and review includedMe, in a session I’m present for
RedResearches it. For red that is not one of the four constitutional acts, runs the reduction ladder belowMe

Ambiguity fails to yellow. Never to green, which is my assurance, and never to red, which would let the system lock itself down by being uncertain.

There is also a fourth disposition, beyond those three, for work that is fully reversible but taste-laden, where staff genuinely cannot infer what I would want: it asks first, in two or three options with a recommendation, capped at two open at a time so that asking cannot become homework.

What makes an item green

An item is green only if all of these hold:

  • it writes to Command’s own files and nowhere else;
  • it lands as a pull request, never a direct push, so every merged change is one revert away;
  • its diff avoids a written danger list (data stores, secrets, anything auth, workflow, schedule or permission related, history rewrites) and stays under caps on both changed lines and changed files, which I don’t reproduce here because they are tuned by signature and would go stale on this page faster than anything else on it;
  • CI is green with a review verdict recorded against the commit being merged.

That verdict is signed with a key the merge gate checks, and the signing requirement is mid-rollout, so today it is policy the gate is growing into rather than a check it fully enforces.

What a stalled yellow hold becomes

A yellow hold older than three days converts itself into a decision on my desk, so a hold becomes a counted interruption instead of an invisible stall.

The four red acts

Four acts are constitutionally red, and no evidence brief or reclassification can ever turn one green.

ActCan I pre-authorize it?
Spend moneyYes, inside an explicit standing bound. I have granted none
Publish or go publicYes, inside an explicit standing bound. I have granted none
Message somebody newNever
Destroy data irreversiblyNever

So all four are per-instance today. Inside a bound, staff acts alone.

The red ladder

For everything else red, staff runs a ladder before it reaches me.

  1. Check whether the label is even correct.
  2. Look for a restructuring that makes the whole thing green.
  3. Carve off the green majority and shrink the red part to its smallest kernel.
  4. Stage that kernel as a one-word decision.

Researching a red item is itself green, so the ladder runs while I’m asleep.

What is enforced by code, not by a rule

Two parts of that are enforced by code, not by a model remembering a rule. The four red acts are a gate function with tests, strictest-wins, that fails closed on anything it cannot evaluate. And a repository Command works on never has its own instruction files loaded: the agent is started with an explicitly empty settings source, so an untrusted repo cannot inject instructions into the session working on it. That second decision is the one that changed the shape of the project. Command stopped consuming per-repo configuration and started authoring it, and the config surface became the product.

The decision folder

This is what the coloring actually routes to me. Every figure in the table carries its source: [measured 2026-08-16, the decision folder’s store].

Decision folder, all timeCount
Items filed27
Signed20
Declined6
Still open1
Resolutions where the decider was me9 of 26
Re-tiered by staff, after re-running the green checklist concluded they had never needed me16 of 26
Recording no decider at all1 of 26

The count is worse than it first reads. All 6 declines are in the re-tiered group, so I have declined nothing. A decline rate computed off that number would be measuring staff correcting its own filing rather than me disagreeing, so the drift signal the rulebook wants from it does not exist yet. Whether the right things reach me at all is the same question from the other side, and the one day anybody measured it, the answer was no.

Promotion and demotion

The system runs a control loop on its own authority.

TriggerWhat happens
A class of decision I have signed clean twelve times runningGets proposed for promotion to green, with the record attached and its first executions still at yellow strength
Any incidentDemotes that class the same day. The demotion sticks to the execution record and not to the paperwork, so signing the demotion does not cancel it
Being asked the same question twiceCounted as a defect, and generates a standing rule instead of a third ask

Execution lanes

Work reaches the code three ways. A shift is one work session clocking in and out against the desk file, and two of the three lanes are kinds of shift.

LaneWhat it isWhat it may do, and what verifies it
Scheduled shiftA shift started by a clock, with nobody presentGreen work only
Attended shiftA session I open myselfThe only tier that can release a yellow hold or stage a red kernel
Commissioned unitA crew a shift hires for a bounded job: a subagent, or a headless Codex run in an isolated worktree with no network accessGreen only. It gets a written packet in and hands artifacts back out, and a shift verifies its output before any of it ships

The routing rule for a commissioned unit is capability, not size: heavy reading, auditing and code archaeology go to Codex, and anything needing a browser, a network call, or a test run stays with Claude because the Codex sandbox has none of those.

Execution records

Four records make execution legible after the fact, and each exists because its absence cost something.

RecordWhat it holdsWhy it exists
Launcher ledgerOne machine-readable row per dispatched childFree-form prose about dispatches meant nothing could answer “which of these actually landed”, and a commercially important job once sat queued behind a click nobody made
WikiProvenance-tagged pages, one per project, venture, decision and concept, with every claim carrying how solidly it is knownIt is the synthesis memory
Decision boardWhatever currently needs me, rendered from a committed specBoth of the hand-built ones carried a freshness stamp that did not match when their facts had actually been read, and one of them showed two finished jobs as still running
Artifact registryThe permanent link of every page a session publishes for meA session that means to update a page but can’t produce its link publishes a second one instead

Being a record is not the same as being right. The first of those four is also in the list of things that don’t work below, because it was wrong on the day I wrote this.

Scheduled jobs

Nine scheduled jobs [measured 2026-08-16, the loop registry]. The load-bearing ones:

JobWhen it runsWhat it does
Morning letterAssembles at 07:00The one thing pushed at me daily
Office shiftNightlyRuns self-checks
Session digestNightlyDigests the day’s sessions into the wiki
Worktree sweepWeeklySweeps for leftover git worktrees and checkouts parked on the wrong branch

Two things make that set safer than it sounds. The first is that an unattended session inherits only the permissions already sitting on disk, so the command set those jobs actually use is committed to the repository rather than clicked through once. The first unanswered permission prompt hangs a run silently, which is exactly what cost the July 24 firing two and three quarter hours. The second is that merging is deliberately not in that set. An unattended shift proposes a merge and records it; it never performs one.

Loop cards

Nothing here ships to a schedule without answering five questions first. Four of them come from control theory, the field behind the thermostat, and the fifth earned its place from experience.

FieldWhat it has to answer
SetpointWhat “correct” means, written so something could check it. A thermostat’s is 20°C
SensorHow the job measures what is currently true. The thermometer. A job that reports its own behavior (“advanced one item”) is self-reporting: it measures the worker, not the world
Error signalThe gap between those two. This is the part that’s usually missing. A count of activity is not an error signal. A backlog depth measured before and after is
ActuatorThe thing that changes reality to close the gap. The furnace. Writing a note for a person to act on later is not an actuator, it’s a sensor with extra steps
Silent failureHow “I could not determine” is distinguishable from “everything is fine”. This is the one that came from experience rather than from the theory

Those five answers are a loop card, and every Command artifact that runs or is trusted unattended owes one, or an exemption from a closed list of six classes. They live in one JSON file, a lint checks the file’s shape and coverage, and npm test runs the lint, so a new workflow or hook with no card is a red build. The scheduled jobs are the exception: they are created in the Claude desktop app rather than as files in the repository, so nothing in a build can see that one exists. The nightly shift re-checks those from the machine itself, after the fact.

Exemptions

The exemptions matter as much as the cards. Three of the six classes:

ClassWhat it covers, and what follows from that
interlockA guard that blocks a bad state. It has no target to converge on, so it owes a liveness answer instead, because nothing else proves a guard is still firing
instrumentA job that reads the outside world and can’t change it. Owes the name of whoever acts on it
open-by-designA loop that ends at a person on purpose. Closing it would be a governance failure, not an engineering fix

Registry state

Every artifact that runs unattended is declared, and every declaration is one or the other: a card, or an exemption. Every figure in the table carries its source: [measured 2026-08-16, the loop registry].

Loop registryCount
Declared artifacts41
Carrying a card22
Declaring an exemption19
Cards recording a known gap22 of 22

A complete card means the job is honestly described, not that it is healthy.

What doesn’t work

The whole design rests on measuring itself honestly, so a page that skipped this part would be advertising the opposite of what the system is for. Every figure below comes from Command’s own operating record and its August control-loop audit, re-read on August 16, not from a separate count taken for this page.

What brokeWhen
The one alarm that lived off this machine never workedAugust 3 to 5
A watchdog on this machine cannot detect this machine stoppingOngoing
Reporting is not acting, and the nightly shift didn’t know the differenceWork half retired August 8
The weekly worktree sweep reported success while the script it calls did not existJuly 27 to August 5
The ledger built to answer “which of these landed” was itself wrongAugust 16, the day I wrote this
The guard that stops a session writing outside its own repo did not fire when it was testedJuly 12
The always-on claim is still unproven, by its own barThe July 24 firing
Every item waiting on me was work staff could already have doneMeasured August 4
The self-improvement loop is built and hasn’t closed yetOngoing. Signal store measured August 16; the directive store behind it last written June 19
The desk file goes stale silentlyAugust 12

The one alarm that lived off this machine never worked. A scheduled cloud job was supposed to file an issue if the morning letter failed to appear. It ran on three mornings, August 3 to 5, and filed nothing on any of them. Its permissions were missing one line, so the second of the two places it looks for the letter could never be read. On two of those mornings it found the letter in the first place it looked and exited green without ever reaching the broken path. The third was August 5, the one morning the letter was genuinely missing, and that is when the broken path ran for the first time: a permission refusal, a not-found, an exit code, and the step that raises the alarm never reached at all. Standing it down removed no working detection, which is the only reason standing it down was cheap. Nothing off-machine has replaced it, so the detector for a missing morning letter is now me noticing.

A watchdog on this machine cannot detect this machine stopping. Every run record these jobs write is local and never pushed, so no off-machine sensor could read one even when one was running.

Reporting is not acting, and for a long time the nightly shift didn’t know the difference. It reported advanced N, a count of items it had touched, and nothing ever re-measured the backlog. So eleven consecutive firings that moved nothing read as eleven separate flags instead of one broken loop. Its work half was retired on August 8 after fourteen firings produced no merged output, stopped by four different blockers. It now runs self-checks and clocks out, and seven of its nine probes only report: a defect they find becomes a flag that may get picked up, and nothing owns closing it or measures the wait. What did get fixed is who does the counting. The backlog is measured at both ends of the run by the always-on service, and the shift is given no way to write that number, which is the arrangement the whole card discipline exists to force.

A job reported success every week while the script it calls did not exist. The weekly worktree sweep was inert from July 27 to August 5 and reported success throughout. Every firing hit a missing file, wrote “script missing” into a report, and exited 0.

The ledger built to answer “which of these landed” was itself wrong on the day I wrote this. That is the launcher ledger from three sections up, the one I credited for making dispatched work legible. This morning’s letter is titled “the dispatch board is under-reporting finished work”, and a commit that afternoon stamped nineteen stale rows terminal and recorded a dispatch that had no row at all. A record that reads as authoritative and lags reality is the same failure as a job that reports its own activity.

The guard that stops a session writing outside its own repo did not fire when it was tested. That was a live cross-runtime test on July 12. It is why the other runtime is capped at proposing rather than releasing, and why the real enforcement is merge-time inspection of every diff, which no runtime can dodge. The hooks are belt-and-braces, and the documentation says so instead of implying a wall.

The always-on claim is still unproven, by its own bar. A scheduled run has never completed cleanly with zero permission prompts and nobody present. The July 24 firing was punctual and did assemble and ship a real letter, then hung about two and three quarter hours waiting on prompts. Until a firing clears that bar, sessions I open myself are the standing fallback, not scaffolding.

On one day in August, every item waiting on me was work staff could already have done. All seven of them, measured on August 4, were green by the checklist. Zero of the seven should have reached me, every one was filed by a job with no authority to release it, and the same item re-filed on four separate days. Nothing measured that, which is why it ran for nine days instead of two.

The self-improvement loop is built and hasn’t closed yet. Command mines recurring friction out of my session history into ranked signals, and before any fix runs, that fix and a success criterion go into git first. Right now the store holds 29 signals: 6 new, 18 corroborated, 5 proposed, and none past that [measured 2026-08-16, the signal store]. So only five have a fix proposed at all, and not one carries a criterion, because not one has been applied. Nothing has gone through end to end. The standing-rule lifecycle has the same shape of hole from the other direction: three rules exist, all three are parked at piloting, and the store has not been written since June 19 [measured 2026-08-16, the directive store]. The measurement machinery around it is the part I trust: every grader allowed to say “proven” has to demonstrate its false-positive rate on synthetic no-effect data first. Three of the runtime paths do; a fourth skips it deliberately and is never allowed to present its output as proof. On that data, naive consecutive-poll comparison false-proves 37 to 42% of the time, and the shipped recipe, which spaces its readings hours apart, false-proves at or under 1.5%. Both are simulation figures, and no real fix has yet been through the recipe.

The desk file goes stale silently. On August 12 the repository merged fourteen pull requests and I made sixteen decisions, and not one session clocked in. The file went on describing the previous day’s world, reading authoritative at a glance, with nothing to indicate its newest row was a day old.

Build the smallest piece of this yourself

The loop card is the part of this that transfers. It needs no company model, no rulebook, and no background service, and it does real work in an afternoon: it makes every job in your repo answer whether it measures anything, or only reports that it ran.

Paste this into Claude Code at the root of a repository that has at least one scheduled job. It offers five exemption classes where Command’s registry has six: the sixth covers a check that is deliberately a fail-open duplicate of an authoritative one, which is a problem you get later.

Find everything in this repo that runs, or is trusted, without a person reading its output each
time: scheduled workflows, git hooks, and any script that decides something on its own. Ordinary
code someone runs and immediately sees the result of does not count.

Give each one an entry in a new `loop-cards.json`, with these five fields:

- `setpoint` - what "working" means for this job, written so something could check it.
- `sensor` - how the job measures what is currently true. A job reporting its own activity
  ("indexed 12 documents") is self-report, not a sensor: it measures the worker, not the world.
- `errorSignal` - the gap between those two. A count of activity is not an error signal. A backlog
  depth measured before and after is.
- `actuator` - what changes reality to close the gap. Writing a line in a log for a human to read
  later is not an actuator.
- `silentFailure` - how "I could not determine" is distinguishable from "everything is fine".

Fill them in from what the job does today, not what it should do. If a job has no actuator, or its
sensor is really self-report, write that in the field. An honest "none" is the useful answer; an
aspirational one makes the file worthless.

Some things fire that trigger and should not get a card. Give those an `exempt` class and a reason
instead of the five fields, choosing from exactly this list:

- `interlock` - a guard that blocks a bad state. No target to converge on.
- `instrument` - reads the outside world; no actuator is available or appropriate.
- `batch-cleanup` - no target state, and visibly wrong next cycle if it did not run.
- `one-shot` - fires once and produces a judgment.
- `open-by-design` - ends at a human on purpose.

Then write `scripts/loop-card-lint.mjs` that:

- finds every scheduled workflow and hook in the repo;
- fails, naming the artifact, when one has no entry in `loop-cards.json`;
- fails when an entry has a missing or empty field, or an exemption class outside that list;
- fails, reporting "detector broken", if it finds zero scheduled workflows, rather than reporting a
  clean repo. A checker that looks at nothing must not be able to return "fine". Print the counts it
  found either way.

Wire it into `npm test` so CI fails when a new unattended job arrives with no card. Add a short
README section defining the five fields and saying the file records what jobs actually do.

Finish by showing me `loop-cards.json` and saying which entries came out worst.

What it produced. I ran that text cold twice on August 16, 2026, each time in a fresh session with no knowledge of Command, against a scratch repository built to contain the defect the exercise looks for: two scheduled workflows, one push-triggered CI workflow, a nightly script that prints a count and exits 0 whatever happens, and a “pre-commit guard” whose entire body is process.exit(0) with no hook installed that would call it.

Both runs found both traps and wrote them down instead of writing something flattering. On the nightly job both filled errorSignal and actuator with None, and used the sensor field to explain that its directory count prints the same number whether the indexing works perfectly or was never implemented. One went further and ran the script: an empty input directory prints “indexed 0 documents” and exits 0, the same shape as a healthy run. Both caught the guard, which refuses nothing, is invoked by nothing, and will be found by anyone grepping for a size check who then believes one is running. Both also noticed the repository’s own CI was a vacuous gate, since node --test on a repo with no test files reports zero tests and exits 0.

Then I tested both checkers myself rather than trusting either run’s account of itself. In both: a new scheduled workflow with no entry fails and names the file; a missing field, an empty field, and an exemption class outside the list each fail; hiding the workflow directory reports detector broken and exits 1 rather than reporting a clean repo; and npm test exits non-zero in every one of those cases, which is the difference between a gate and a report. About ten minutes each, and neither needed a correction, so what’s above is the first version and not a third draft. Both runs found both traps in a repo I had salted with them, though. Whether it finds the ones in a repo nobody prepared is the part two runs cannot establish.

Where the two runs differed, which is the real limitation. The prompt doesn’t pin the file’s shape, and they chose different ones: one keyed entries by path inside an object, the other used a list with the path as a field. Each run’s checker matches its own file, so both work, but if you want a particular schema you have to ask for it. Both runs also added a failure the prompt never asked for, when an entry points at a file that no longer exists, and I checked that both really do it. That is the right call and I would have missed it: without that check the file rots quietly, which is the failure the file exists to prevent.

What this does not give you. It won’t tell you whether a card is honest. The lint can check that the sensor field is filled in; it cannot tell a real measurement from a self-report dressed up as one, and that judgment is the whole value of the exercise. Command has exactly the same hole, written down as deliberately cut rather than quietly missing: the fix is a review lens that reads cards adversarially, and it isn’t built. The check is also narrower than the inventory that precedes it. It re-detects scheduled workflows and hooks, so a new one of those fails the build, but a plain script that quietly starts deciding something on its own is not mechanically detectable as such, and neither run pretended otherwise. Both caught the guard script in my scratch repo only because its own header said it was a hook. And the check sees nothing outside the repository at all, which for Command is 9 of its 41 declared artifacts. Treat the file as the place the question gets asked, not as proof that anybody answered it well.

Last verified August 16, 2026. That is when the prompt was last run cold and the checker it produced was re-tested by hand. If that date has aged by the time you read this, treat the prompt as unverified, because that is exactly what it would be.