Skip to content
CrewibleGet Crewible
← The whole crew

Audit

Judy

Judy is the auditor. Every other crew member builds, writes, ships or reports; Judy checks whether what the system says about itself is actually true, using evidence rather than assertion. Read-only, escalate-only, and it runs after the day's work, never inside it. The one rule: a claim is not evidence.

I built Judy after one afternoon turned up four faults with the same shape. A config said an integration wasn't connected; the key was present and working, and 'connect it' was ranked the top action for that brand. The money agent had reported 'nothing to report' for as long as the system existed, because a permissions list had quietly left out the one tool it needed, and a truthful nothing is indistinguishable from an empty ledger. Nothing errored. Everything was cheap to detect. Nothing had the job of doubting the system, so nothing did.

That job is now Judy's. It takes the claims the crew makes about itself, a connector is live, an escalation was sent, there's nothing to report, and asks what each one rests on. A record that shows something actually happened is evidence; a status file saying so is just a claim. Where it can't get evidence, it says 'I could not verify this' in those words, because 'probably fine' is exactly the reasoning it exists to catch. Verdicts are ranked by what a wrong claim cost you, not by how dramatic it sounds.

What Judy does

  • Audits the crew's claims about itself against evidence: records, logs, outputs. Not vibes.
  • Hunts the quiet failures: the agent that reports nothing every day, the solved blocker still ranked urgent, the promise kept nowhere.
  • Returns short rulings, evidence first: contempt (false and acted on), suspect (unsupported), note (true and worth knowing).
  • Says 'I could not verify this' when it can't, and 'sound today' in one line when it is.

One audit, start to finish

The day's work is done and the crew has filed its reports. Judy's session starts after all of it, on purpose.

  1. 1

    Judy gathers what can be checked mechanically first: which claims have records behind them and which are just statements in a file.

  2. 2

    It looks for the classic: an agent that has reported 'nothing' every single run. Before believing there's nothing, it checks the agent could actually act. That's how a missing permission hides for months.

  3. 3

    It reads the ranked lists: is anything still flagged as blocking that the records show was solved weeks ago? Stale blockers send real effort to solved problems.

  4. 4

    It checks promises against mechanisms: the vault says the business never does X. Is there anything that would actually catch X, or does the rule rely on everyone remembering?

  5. 5

    It rules: evidence first, verdict second. Contempt for the false claim that was acted on, suspect for the unsupported one, note for what's true and worth knowing.

  6. 6

    The ruling lands with you, ranked by what each item cost. If the system is sound, the entire output is one line saying so. A quiet court is a real result.

What a real Judy ruling looks like

JUDY · Daily audit
CONTEMPT (1): config says the payment connector is not set
up; the key exists and has processed live charges. "Connect
payments" is ranked your #1 action. It's done. Re-rank.
SUSPECT (1): Finn has reported "nothing to report" 11 runs
straight. Verify Finn can actually reach the data before
believing there's nothing in it.
NOTE: could not verify last week's email escalation was
delivered; no delivery record exists to check. That gap is
itself the finding.

Illustrative example. Every line is a claim checked against a record, which is the entire job.

Where Judy fits in the crew

Bench reviews a piece of work before it ships; Judy audits the system that produced it, after the fact and from outside. It shares Finn's constitution: read-only for the same reason an auditor doesn't edit the books. If a fix is obvious, Judy names it and who owns it, then leaves it.

The honest bit

Judy fixes nothing, ever. Not a config typo, not a stale status, not the obvious one-liner, because an auditor who edits is auditing her own edits. It runs in its own session after the day's work, never inside it. And it will not manufacture findings to justify existing: if the system is sound today, the output is one line saying so, which is the exact discipline it exists to enforce.

Straight from Judy’s actual file

Every crew member is a file you can read and edit. This is a real slice of the one you get, not a mock-up. The full file, and the other eleven, come with the download.

crew/agents/judy.md
## The one rule
**A claim is not evidence.**

A config saying a connector is live is a claim. A record showing it carried
something is evidence. An agent reporting "nothing to report" is a claim, and
the evidence is whether it had a tool it was permitted to run.

When you cannot get evidence, say so in those words. "I could not verify this"
is a finding; "this is probably fine" is not.

Judy comes with the other eleven.

One download, twelve specialists, the brain vault and the honest autonomy layer. Reskin the lot to your business.