How claims are governed here A reenactment of the starter kit's governance loop: the runnable kit that distills this program's claim-governance machinery, downloadable on this site at starter-kit/. Eight stations walk a claim from registration to the ledger, each naming a real file in the kit. The moving tokens are claims quoted verbatim from the 44 assessment packets, each cited below with its packet file and line. The machine does not decide verdicts. It reads the ones the packets already recorded. Said plainly: the 44 verdicts were produced under the assessment protocol in RULEBOOK.md, not by running this kit. The kit came after the program, to make the machinery portable.

Seven definitions before the machine. A claim is a checkable assertion a project makes in public (example on the belt: "Main CI passes at the pinned commit"). CI is continuous integration: the service that automatically runs a project's tests on each commit. "CI green at the pin" means those tests visibly passed on the project's public CI service at the assessed commit. The "pin" is the single commit the assessor fixed for the whole evaluation. NODUS is the program's four-ring verdict scale (Invest, Pilot, Explore, Monitor) assigned per assessment packet. TRL is technology readiness level, scored 1 to 9 per packet. The letter codes (A3, B5, D1) are item ids in the kit's checklist: A = Phase A planning, B = Phase B execution, D = the demotion rules. Every code on this page is defined in starter-kit/CHECKLIST.md. The license rider is the MIT license plus an AI-lab rider: it bars OpenAI and Anthropic from training on the code, using it, or benchmarking it. This page depicts the kit's loop, inspectable and runnable from starter-kit/. The 44 verdicts were made under RULEBOOK.md; the two pipelines are kept separate on purpose (see "Two different pipelines" below).
Paraphrase: the maintainer-writing brief reads Emanuel's Apr 2025 thesis ("Protecting Against AI Prompt Injection") as arguing that internal guardrails always fail, and proposing external "inoculation" strategies instead. The brief quotes his Jun 2024 essay's alternative as "a system to closely monitor and control the AI that is fundamentally external to the AI and not controlled by it", with helper models required to "supply the hard evidence". Paraphrase and quoted lines via synthesis/briefs/maintainer-writing.md (the Apr 2025 thesis is paraphrased there, not quoted; the Jun 2024 lines are quoted there). The older claim that this page once carried, a verbatim Emanuel quote dated 2024 saying "internal guardrails fail", had no source and was removed.
44 repos assessed 42 of 44 couldn't show public CI green at the assessed pin 38 of 44 carry the license rider 0 of 44 independently validated
The assessment protocol behind every verdict on this site ships with it: RULEBOOK.md — evidence tiers, the claim-inventory rule, NODUS ring definitions, and the QA checklist each packet had to clear. The verdicts section below summarizes it.
0registered 0enforced 0challenged 0demoted 0ledger rows
1 · REGISTRY — A4 a claim is logged before it is boasted templates/claims.tsv · one row per planned claim, proof slot included 2 · READINESS — A3 READY · 12/12 sections scripts/check-readiness.sh · the plan, once, at the gate 12 sections · structure, not truth · A13 is the semantic backstop 3 · HOOK + CANARY — B5 ! canary rejected: teeth confirmed the pre-commit hook · claim check + ledger lint, every commit first it must reject a planted overclaim (scripts/hooks/pre-commit) 4 · README BOARD — B6 README.md enforced claims pin here, each pointing at a proof artifact a public README with zero enforced claims fails the check 5 · CI BACKSTOP — B5 escape hatch templates/kit-gates.yml re-runs every gate where --no-verify can't reach the kit's workflow. None of the 44 repos runs it 6 · REVIEW — A6 · A13 · B7 · B8 ? ? ? the packet's verdict, read aloud A13: independent review is the semantic backstop demotion rules run on schedule by review, never automatically 7 · DEMOTION — D1–D7 D1 · D4 · P0 tombstone demotion needs counter-evidence. DEMOTED promotion needs a gate; demotion needs counter-evidence rules written down in advance (A6) 8 · LEDGER — A5 · B4 rows banked: 0 templates/negative-evidence-entry.md · row schema the hook lints every staged row. Weasel predicates block the commit Token labels are truncated prefixes. The full claim text is quoted in the claims panel below, with packet file and line.

The loop in words

  1. Registry: the claim is logged in templates/claims.tsv (installed by the kit as registries/claims.tsv) before the README gets a sentence. (A4)
  2. Readiness: scripts/check-readiness.sh machine-checks the plan once at the phase gate: 12 sections, structure not truth. (A3)
  3. Hook: the pre-commit hook runs the claim check, and first fails a planted canary to prove it bites. (B5)
  4. README board: enforced claims pin to the board, each pointing at a proof artifact; scripts/check-claim-discipline.sh does the cross-checking. (B6)
  5. CI backstop: templates/kit-gates.yml re-runs the gates where --no-verify can't reach. The kit's workflow, not the repos'. (B5)
  6. Review: independent review is the semantic backstop; the demotion rules are written down in advance (A6) and executed on schedule by review (B8), never automatically. (A6, A13, B7, B8)
  7. Demotion: counter-evidence arrives, the rules fire, the stamp lands. (D1–D7 in demotion-rules.md)
  8. Ledger: hypothesis, verdict, retry predicate: banked in the row schema of templates/negative-evidence-entry.md, never "later". (A5, B4)

How the verdicts were made

The briefs on this site were produced under the assessment protocol in RULEBOOK.md. What follows is that document's own wording, condensed. The full protocol — including the QA checklist every packet had to clear — ships with the site.

Section map. This page publishes four Rulebook sections: §0 item 3 only (quoted under Evidence tiers below), §1 (the tier table), §4.3 (claim-inventory statuses, on the conveyor), §4.9 (ring assignment rules, including the close). Everything else in the protocol (§2 source hierarchy, §3 pinning and cutoff, the rest of the §4 packet template, §5 deepening questions, §6 cross-cutting lenses, §7 language rules, §8 QA checklist) lives only in RULEBOOK.md, which ships with the site.

NODUS — the four rings

NODUS is the program's verdict system: each packet assigns one of four rings, naming what the evidence supports doing with the repo. Justified in one line per criterion (technology readiness, strategic relevance, impact potential, implementation feasibility, time to mainstream, collaboration potential). Ring assignment rules, Rulebook §4.9:

Invest0 of 44

Requires independent validation plus governance. No repo in the suite holds it: 0 of 44 are independently validated (synthesis/00-overview.md).

Pilot3 of 44

Requires a release artifact plus a bounded, real workload fit. The packets place 3 repos here: asupersync, franken_agent_detection, franken_ocr. "Pilot" names the fit; it is not a recommendation to ship.

Explore34 of 44

The default for substantive-but-unproven. Most of the suite lives here.

Monitor7 of 44

Websites, retired artifacts, and plan-stage work. The four website repos, the two beads dashboards, and franken_nlp (plan-stage).

Rulebook §4.9, last line: "When in doubt, ring down, not up."

TRL: technology readiness, 1–9

Each packet scores technology readiness on the 1–9 scale and justifies the score in one line. Observed across the suite: TRL 2–3 up to 9 (the 9 is one packet's own rating of a deployed website artifact, not a suite-wide norm). TRL measures readiness, not quality.

Evidence tiers — every claim carries one

Rulebook §0: "Every substantive claim carries an evidence tier and a confidence grade. A claim without both is a draft note, not a finding." Confidence grades: High (multiple converging sources or direct inspection), Medium (single solid source, plausible), Low (thin evidence, extrapolation).

1 · VerifiedConfirmed by direct inspection of a fresh clone, an API response, or a live page read by the analyst.
2 · CI-observedObserved executing on live CI pages. Attests the suite runs, not that it is green, unless pass/fail is legible.
3 · Maintainer claimAsserted in README/docs by the maintainer; not independently executed or reproduced.
4 · ExternalIndependent sources: APIs, papers, press, third-party benchmarks. Absence of coverage is reported as a finding.
5 · InferenceThe analyst's judgment. Always labeled; never presented as fact.

Checklist items outside the diagram

The animation walks eight enforcement points. The kit's checklist (starter-kit/CHECKLIST.md) has 28 items. Six more that matter, in the checklist's own terms:

Independent review — the semantic backstopA13

A second agent (fresh eyes, different model, or adversarial duel) reviews the packet and records what changed. The machine verifies the attestation exists; it can never verify the review's quality. READY is not the same as reviewed.

Anti-reward-hacking lawA9

The 12 forbidden reward-hacking patterns are named verbatim in the project's agent instructions: gate self-weakening, proof-class inflation, golden regeneration reflex, commit-stream pumping, tautological tests, easy-lever cherry-picking, close-pump abuse, scope-splitting, spec-editing as progress, conformance metastasis, dependency smuggling, bench-path hardcoding — under three load-bearing rules: never weaken a gate to land a change; no self-grading without independent verification; demotions are always allowed. Verified by review.

The beads graph is the executable formB1

The beads graph is the plan of record; documents explain why, beads say what is being done and whether it closed. .beads/issues.jsonl is non-empty and current; no parallel prose tracker exists.

Close only on cited evidenceB2

A bead closes only with a close_reason citing the evidence (commit, receipt, ledger row); prose assertions never close beads. Verified by review of closed beads.

Waivers are public, time-bounded, recordedB9

Any waiver of a release-blocking gate states rationale, owner, expiry, and compensating controls; waivers bypassing a release-blocking gate force NO_GO unless all four are recorded and signed off. No waiver outlives its expiry.

Session completion: land the planeB10

Every work session ends with the landing ritual: file beads for remaining work, run the quality gates, update bead states, sync the tracker, hand off context.

Two different pipelines, kept separate on purpose

The CI backstop in the diagram is the starter kit's workflow, shipped as templates/kit-gates.yml (installed by init.sh as .github/workflows/kit-gates.yml). Its job is to re-run the gates where git's --no-verify can't reach. None of the 44 assessed repos runs it.

What the 44 packets found about the repos' own CI (synthesis/00-overview.md. These are assessment findings, not this machine's output): public CI green at the pin, 2 of 44 (frankenscipy, franken_threed) · red at the pin, 11 · CI exists but no pin verdict, 6 · no test CI / deploy-only, 10 · CI disabled or deleted, 8 · private-only and unobservable, 7. The 42-of-44 figure measures whether public CI was green at the assessed commit. It measures CI, not claim demotions. Disclosure: frankenjax is filed C4 in the overview matrix; this site reclassifies it C6 per the legend's RCH criterion (packet: all green claims execute on the maintainer's private RCH fleet, attested only by checked-in JSON).

Legend — what each part is, in the program's own terms

Claims registryA4

Every public claim gets a row in templates/claims.tsv before the README gets a sentence. It carries a status and a proof slot. A claim registered after the fact is a rationalization, not a claim. (The suite's best coverage audits found registries capture only 2.0% and 6.8% of claims in the wild, checklist item B12. The registry is the floor, not the ceiling.)

Readiness gateA3

scripts/check-readiness.sh machine-checks the plan once at the phase gate: 12 sections, and a missing one means NOT READY. It checks structure, not truth. Independent review (A13) is the semantic backstop.

Pre-commit hook + canaryB5

The pre-commit hook runs the claim check on every commit, and first fails a deliberately overclaimed canary to prove it still has teeth. A gate that can't fail is decoration. (git's --no-verify bypasses the hook silently; the CI backstop is the answer to that.)

The README (enforced claims)B6

Claim discipline: every enforced README claim must resolve to a real proof artifact, cross-checked by scripts/check-claim-discipline.sh. A public README with zero enforced claims fails the check. README drift is a governance failure, not documentation debt.

CI backstopB5 escape-hatch clause

templates/kit-gates.yml re-runs every gate where --no-verify can't reach. This is the kit's workflow, installed by init.sh. The 44 repos' own CI results are assessed separately (see the panel above).

ReviewA6 · A13 · B7 · B8

Independent review (A13) is the semantic backstop. The demotion rules (D1–D7) are written down in advance (A6) as templates/demotion-rules.md and executed on schedule by review (B8): expired manual proofs demote, open P0s block release claims, retired ids are tombstoned. Gate changes need two-direction evidence (B7). All procedural: nothing here is automatic.

Demotion

Promotion needs a gate; demotion needs counter-evidence. The starter rules live in templates/demotion-rules.md. In this reenactment the counter-evidence is the assessment packet's own recorded verdict, quoted in the claims panel, cited by file and line.

Negative-evidence ledgerA5 · B4

The project's memory of what didn't work: hypothesis, verdict, retry predicate, lesson (row schema: templates/negative-evidence-entry.md). The pre-commit hook lints every staged row and rejects weasel retry predicates ("later", "TBD", "n/a"). Every reject needs a retry predicate. Never "later".

The claims on the conveyor — quoted verbatim

Every token carries one of these claims, quoted word-for-word from the claim inventory (Rulebook §4.3) of its assessment packet. Each entry names the packet file and line, the packet's recorded status, and — for demoted claims — the packet's own counter-evidence. The twelve claim strings are verbatim from their packets. The counter-evidence blocks are each packet's own recorded verdict, lightly normalized for display (stripped Markdown, quote-style changes, truncation of trailing detail) — the quoted wording is the packet's wording, and one block adds a one-sentence plain-language gloss outside the quotation marks.

franken_threedfranken_threed-assessment.md:94packet: demonstrated · HELD
"Differential harness: all 256 cube cases × 8 flag combos + animated metaballs + 80 seeded fields + NaN/empty edge cases, asserted against the live upstream oracle"
No counter-evidence in the packet. The assessor independently re-ran the suite: 61/61 pass. The claim survives review.
frankenlibcfrankenlibc-assessment.md:86packet: Disproven · DEMOTED
"Main CI passes at the pinned commit"
Counter-evidence (packet): "Pin's main CI run #8942: Core Gates failed on its first step (`scripts/ci.sh`); all substantive jobs skipped. A separate candidate-validation workflow for the pin's change was green, but that is not main CI — [High]"
franken_numpyfranken_numpy-assessment.md:102packet: demonstrated · HELD
"Negative-evidence ledger: 67,641 lines of recorded losses, no-ships, reverts"
No counter-evidence in the packet (counted at the pin). The claim survives review.
frankenlibcfrankenlibc-assessment.md:93packet: Disproven · DEMOTED
"Runtime depends on `asupersync-conformance`"
Counter-evidence (packet): "Verified: dev-dependency in membrane, optional/feature-gated in fixture/harness/conformance."
franken_markdownfranken_markdown-assessment.md:70packet: Demonstrated · HELD
"Engine library has zero third-party dependencies"
No counter-evidence in the packet. The claim survives review.
frankenredisfrankenredis-assessment.md:89packet: partially demonstrated · DEMOTED
"Full upstream Redis 7.2.4 Tcl suite runs in CI (284 runs observed)"
Counter-evidence (packet): "Greenness at the pin: RED — scheduled run #284 (head = pin) executed the suite (~21 min) but the 'Verify complete upstream suite verdict' step failed." The suite ran; green is what failed.
frankenfsfrankenfs-assessment.md:75packet: Demonstrated · HELD
"No `todo!()`/`unimplemented!()` in production code — and the ban is self-enforced"
No counter-evidence in the packet. The claim survives review.
franken_markdownfranken_markdown-assessment.md:72packet: Aspirational / unmeasured · DEMOTED
"Ultra-fast" (crate + GitHub description)
Counter-evidence (packet): "[Verified absence, High] — no throughput numbers (no MiB/s, no head-to-head vs pulldown-cmark/comrak) anywhere in README/docs."
frankenredisfrankenredis-assessment.md:96packet: demonstrated · HELD
"NEGATIVE_EVIDENCE.md: 26,485 lines of recorded losses with verdicts"
No counter-evidence in the packet (counted at the pin). The claim survives review.
franken_markdownfranken_markdown-assessment.md:77packet: Stale wording; mechanism real · DEMOTED
"Every marketed claim CI-enforced ("CI fails if it drops")"
Counter-evidence (packet): "The only GitHub Actions workflow is named 'DISABLED — releases use DSR exclusively' with `if: ${{ false }}`; the README's 'CI' means 'the maintainer's machine,' not CI." The gate scripts are real; the word "CI" is stale.
frankenredisfrankenredis-assessment.md:92packet: demonstrated · HELD
"Those benchmarks are disavowed pending a deferred re-baseline bead (`vibu6`)"
A disavowal, documented: the README itself says the numbers "should be read as a contention-sandbox measurement, not a release claim". The claim survives review.
frankensearchfrankensearch-assessment.md:89packet: demonstrated · HELD
"Two-tier progressive search: Phase 1 `Initial` delivered sub-ms; Phase 2 `Refined` in ~6 ms"
No counter-evidence in the packet. The claim survives review.

Ledger rows — what demotion leaves behind

Nothing banked yet. A demoted claim's obituary lands here: the claim, the packet's verdict, and the counter-evidence it was quoted from.

Machine log