# FrankenPandas — RULEBOOK v1.0 Assessment Packet v5

**Repository:** `Dicklesworthstone/frankenpandas` · **Language:** Rust (edition 2024) [Code-verified, High] · **Pinned commit:** `ebe79a3b4ae0470402ea76f9547ae93093360b37` (2026-09-22 13:13:20 UTC, [Git-observed, High]) · **Last push:** 2026-09-22 (pin date; the pin is the latest push confirmed in this assessment) [Git-observed, High] · **Scope:** the pinned commit only, not HEAD. No tag points at the pin; the newest tags are the `v0.3.0` / per-crate `fp-*-v0.3.0` release tags (2026-09-12) pointing at earlier commit `0ad22f91` [External, High]; ten GitHub Releases (workspace + 9 per-crate, all `v0.3.0`, published 2026-09-12, not draft/prerelease) target that earlier tag — no release artifact for the assessed commit [External, High]. *Cold-reader note: "v5" = rewrite-loop round 5 (final). Revision history: v2 added the certification-funnel aggregation, rider steelman (bead `frankenpandas-dio8`), worked groupby trace, and verified `fixture_provenance`; v3 softened the rider-intent certainty (deliberate-retention vs retention-by-default) and added the funnel-flattery mechanism check; v4 was verification-only (counts re-run, links re-fetched, method note corrected). "v1" = the from-scratch draft.*

**Method (analyst):** shallow *sparse* checkout at the pin under `~/workspace/.scratch/frankenpandas-verify` (41 MB working tree: `crates/`, `scripts/`, `docs/`, `.github/`, `.beads/`, `artifacts/perf/`, root manifests — the full `git clone` and a 350 MB tarball both failed on this network; `artifacts/bench` (960 rows, 31 MB) was added to the sparse checkout in a later pass, but `artifacts/phase2c` and the fuzz corpus were not pulled — findings that need them are marked accordingly). Read: root `Cargo.toml` (15 members enumerated), all 15 crate manifests, `LICENSE` (verbatim, plus the raw fetch), README (2,976 lines: drift, disclaimers, heat map, contributions policy), `CHANGELOG.md`, `crates/fp-conformance/DISCREPANCIES.md` (28 entries), `docs/NEGATIVE_EVIDENCE.md` (sampled: methodology rows 112/116/118), `artifacts/perf/SCORECARD.md` (generated 2026-09-03), `benches/BENCH_MATRIX_SPEC.md` (existence), `scripts/current_loss_census.py` + `scripts/validate_api_coverage_drift.py` (headers), `crates/fp-runtime/src/lib.rs` (RuntimePolicy/RuntimeMode/RaptorQEnvelope/ConformalGuard/asupersync gating), `crates/fp-dot-kernel/src/lib.rs` (ISA-isolation docs), `deny.toml` (advisory ignores), `.beads/beads.base.jsonl` (4,051 records). Counted: `.rs` files/lines under `crates/` and under `src/`, `forbid(unsafe_code)` per crate root, `unsafe` token occurrences outside comments, `#[test]` attributes, packet JSONs, DISC entries, beads by status. **Not done:** the workspace was never compiled, nothing was executed, no test suite was run, no benchmark was reproduced, no crates.io lookup succeeded (crates.io API returns HTTP 403 from this network — publication status unverified), CI job logs need admin rights (failing-step detail beyond annotations unestablished), line-level authorship attribution not performed. Assessment date: 2026-09-22.

**Tier legend (Rulebook §1):** **[Verified]** direct inspection of the pinned sparse checkout or a live page read by the analyst — with flavors **[Counted]** (I ran the count), **[Git-observed]** (git/API metadata), **[Code-verified]** (source read); **[CI-observed]** is Tier 2 (seen executing on live CI pages — attests the suite *runs*, not that it is green); **[Maintainer claim]** asserted in README/docs, not independently executed; **[External]** independent sources; **[Inference]** analyst judgment, always labeled. Confidence: **High** (multiple converging sources or direct inspection) / **Medium** (single solid source, plausible) / **Low** (thin evidence, extrapolation).

---

## Hook

A 648,671-line [Counted, High], zero-`unsafe` [Counted, High], single-maintainer Rust reimplementation of the pandas API that beats pandas 2.2.3 by a certified 3.97x geometric mean across 143 of 359 measured benchmark lanes [Code-verified (generated scorecard), Medium] — and keeps a 44,086-line public ledger of every time its own numbers were wrong, including the entry where a claimed 1.20x win turned out to be build variance [Code-verified, High]. The catch is threefold: 216 lanes never cleared the certification gate (74 undecidable, 62 dropped on variance — §4.5) [Code-verified, High], its license names OpenAI and Anthropic as forbidden parties and orders destroyed copies on breach [Code-verified (license text), High], and its README explicitly refuses outside contributions [Code-verified (README text), High]. The sharpest irony is saved for §4.8: the older bulk of the tree is agent-co-authored — overwhelmingly by Claude-family models from Anthropic, a named Restricted Party [Git-observed, High on the counts; the implication is Inference, Medium].

---

## TL;DR

- **What it is:** From-scratch Rust reimplementation of the pandas API (648,671 lines / 792 files, 15 crates [Counted, High]), chasing behavioral parity with pandas 2.2.3 — not a binding, not a fork. One human maintainer (Jeffrey Emanuel) plus agent co-authors; 8,621 commits [External, High]; 4,051 beads, 4,049 closed [Counted, High].
- **Strongest evidence:** 1,387 conformance packets counted in-tree (exactly the README's number); 8,954 `#[test]` functions; a 44,086-line negative-evidence ledger with inline perf-claim retractions [all Counted, High]; a generated vs-pandas scorecard with a documented certification gate (live incumbent same-invocation, per-arm A/A nulls, bootstrap median-CI, ELF SHA pinning) [Code-verified, High]; a genuine zero-`unsafe` census — all 15 crate roots `forbid(unsafe_code)` [Counted, High]; 10 certified losses listed with thread counts [Code-verified, High].
- **Strongest doubts:** The 3.97x geomean covers *certified* lanes only — 143 of 359 (216 uncertified: 74 undecidable, 62 high-CV, 78 read-but-uncertified) [Code-verified, High]; the maintainer's own ledger declares sub-1.5x ratios UNRESOLVED after measuring 2.6x same-source build variance [Code-verified, High]; the live-oracle conformance job and `--all-features` lint are red [CI-observed, High]; the license rider bars OpenAI/Anthropic (and anyone acting for them) from use/benchmarking/analysis and demands destroyed copies on breach [Code-verified, High]; the README refuses outside contributions [Code-verified, High]; the 119/119 and 1484/1484 Python-parity numerals are README-only [Maintainer claim, Medium]; crates.io publication unverifiable from this network [Not verified].
- **NODUS ring: Explore** [Inference, Medium] (TRL 4–5; §4.9). Substantive-but-unproven; the rider + no-contributions policy form a double ceiling — the rider a *durable* one (adopted at inception, survived review in bead `frankenpandas-dio8`; §4.8 adjudicates intent at [Inference, Medium]). Track the *methodology* (negative-evidence ledgers, certified-lane benchmarking, differential conformance, fail-closed runtime policy) as a FrankenSuite exemplar; the software is un-advanceable while the rider stands.

---

## Quick Links

Pin-relative links (`blob/ebe79a3b4ae0470402ea76f9547ae93093360b37`); all verified resolving (HTTP 200) on 2026-09-22 [Verified, High]:

1. [Repository](https://github.com/Dicklesworthstone/frankenpandas)
2. [README](https://github.com/Dicklesworthstone/frankenpandas/blob/ebe79a3b4ae0470402ea76f9547ae93093360b37/README.md) — 2,976 lines
3. [License (with AI-lab rider)](https://github.com/Dicklesworthstone/frankenpandas/blob/ebe79a3b4ae0470402ea76f9547ae93093360b37/LICENSE)
4. [Changelog](https://github.com/Dicklesworthstone/frankenpandas/blob/ebe79a3b4ae0470402ea76f9547ae93093360b37/CHANGELOG.md) — 0.3.0 / Phase 4
5. [Known conformance divergences (DISCREPANCIES.md)](https://github.com/Dicklesworthstone/frankenpandas/blob/ebe79a3b4ae0470402ea76f9547ae93093360b37/crates/fp-conformance/DISCREPANCIES.md) — 28 numbered entries
6. [Negative-evidence ledger](https://github.com/Dicklesworthstone/frankenpandas/blob/ebe79a3b4ae0470402ea76f9547ae93093360b37/docs/NEGATIVE_EVIDENCE.md) — 44,086 lines at the pin
7. [vs-pandas scorecard (generated)](https://github.com/Dicklesworthstone/frankenpandas/blob/ebe79a3b4ae0470402ea76f9547ae93093360b37/artifacts/perf/SCORECARD.md) — 143 certified lanes
8. [Benchmark matrix spec](https://github.com/Dicklesworthstone/frankenpandas/blob/ebe79a3b4ae0470402ea76f9547ae93093360b37/benches/BENCH_MATRIX_SPEC.md)
9. [Agent guidelines (AGENTS.md)](https://github.com/Dicklesworthstone/frankenpandas/blob/ebe79a3b4ae0470402ea76f9547ae93093360b37/AGENTS.md) — the project's own anti-reward-hacking rules
10. [CI workflow](https://github.com/Dicklesworthstone/frankenpandas/actions/workflows/ci.yml) — latest run: conformance job red, 15/19 jobs green
11. [Fast gate workflow](https://github.com/Dicklesworthstone/frankenpandas/actions/workflows/fast-gate.yml) — green at the latest run
12. [Fuzz Nightly workflow](https://github.com/Dicklesworthstone/frankenpandas/actions/workflows/fuzz-nightly.yml) — red on the last 5 runs
13. [Releases](https://github.com/Dicklesworthstone/frankenpandas/releases) — v0.3.0 workspace + 9 per-crate releases (2026-09-12)

---

## Did You Know

The SIMD story is the project in miniature. `fp-dot-kernel` is the one crate compiled with `+avx2,+fma` — ISA-isolated so the rest of the workspace stays at baseline — and its module docs record that an `+avx2,+fma` build of the identical source once came out **bit-identical** to the baseline build (both ELFs checksummed `4957dea0fe3e2ed1`): the flags bought nothing, and the real win was reordering the loop from DOT order to AXPY order (0.59 GFLOP/s for a 1000×1000 materialize, cache-miss-bound). The docs also record a measured trap: per-crate rustflags are silently lost if the kernel gets inlined or goes generic in the caller — "a green build with correct values and no speedup, invisible without disassembly" — so the entry point is `#[inline(never)]` and non-generic, both load-bearing. [Code-verified, High — `crates/fp-dot-kernel/src/lib.rs` at the pin.] A project that documents the optimization that *didn't work*, with checksums, is practicing the methodology it preaches.

---

## Franken-worthy next steps

1. **Extract the certified-lane benchmark discipline as the program's evidence template.** The live-incumbent-in-the-same-invocation + per-arm A/A null + bootstrap median-CI gate + ELF SHA pinning + STALE-ELF re-measure flags + "certified losses listed with thread counts" is the strongest benchmark-honesty apparatus in the FrankenSuite — stronger than frankenredis's disavowed tables. *Done when:* the schema is extracted into the program's assessment protocol as the required format for any cited benchmark, with a named owner. [Inference, High — process proposal]
2. **Adopt the negative-evidence ledger with PERF-CLAIM CORRECTION rows.** FrankenPandas' `docs/NEGATIVE_EVIDENCE.md` (44,086 lines) doesn't just log losses — it logs *retractions*, e.g. the 1.20x "win" re-diagnosed as build variance in the same row. Require every program-published performance claim to carry this correction discipline. *Done when:* the requirement is written into the protocol with a named owner. [Inference, High]
3. **Write the supply-chain screen the rider forces.** Any FrankenSuite dependency candidate with a named-party restriction fails intake automatically — frankenpandas is the second case (after frankenredis) that forces it, and its destroy-on-breach clause is harsher. *Done when:* the intake rule is committed to the program's dependency policy. [Inference, High]
4. **Use the asupersync optional-integration as the program's dependency-policy case study.** frankenpandas integrates asupersync as an *optional, off-by-default* feature — and v2 found the integration is deeper than v1 reported: a 1,165-line artifact-transfer subsystem (codec, capability gates, integrity verifiers, recovery policies, transport) inside `fp-runtime`, consumed by `fp-conformance` behind the feature gate, plus an `Outcome` → `DecisionAction` mapping for the runtime policy — while frankenredis evaluated and rejected it. The two postures bound the policy space (adopt-as-optional-subsystem vs. reject-with-reasons). *Done when:* the policy documents both postures with these two repos as the worked examples. [Inference, High]
5. **Falsify the thread-spawn causal story after the shared pool lands.** The maintainer's causal account of the 100k-row losses is precise: per-call `thread::scope` fan-out costs ~350–420µs, which is why several 100k-row lanes lose while the same lanes win at 1M rows; the tracked fix is replacing 147 fan-out sites with one shared pool. *Done when:* the 10 certified losses are re-measured after the pool lands and the story is confirmed or rejected per lane. [Inference, High]

---

## 4.1 Why it exists — the market problem

**The problem, as the maintainer frames it:** pandas is the lingua franca of data analysis, but it is single-threaded Python with unpredictable memory spikes, GIL contention in production pipelines, and dtype-coercion surprises that silently corrupt results — while drop-in *performance* replacements (Polars, DuckDB) require rewriting code in a different API [Maintainer claim, Medium — the README's TL;DR framing]. The maintainer's stated bet is that the entire pandas API can be re-derived from first principles in safe Rust — same semantics, same method names, same edge-case behavior — with columnar storage, vectorized kernels, arena-backed execution, and an explicit alignment-planning phase, verified differentially against the actual pandas oracle on every commit [Maintainer claim, Medium].

**Who feels the pain:** Rust shops that want pandas semantics (index alignment, identical method names, identical edge-case behavior) without a Python interpreter, and Python shops that want a faster `import pandas` — the `fp-python` PyO3 crate is the explicit bridge (`import frankenpandas as pd`) [Maintainer claim, Medium].

**Why now:** the Rust ecosystem (portable SIMD, PyO3/maturin, Arrow/Parquet crates) is mature enough to attempt API-complete re-derivation, and pandas 2.x's own behavioral surface is stable enough to pin an oracle against (pandas 2.2.3 throughout) [Inference, Medium]. Whether that window is real demand or maintainer narrative is unproven — there are no production users, no independent coverage, and the assessed commit has no release artifact [External, High within recall caveats; Inference, Medium].

**Why a rewrite, not a binding:** stated in the "What does clean-room mean" FAQ — behavior is studied via the conformance oracle (input → output contracts, edge cases, dtype rules), never the pandas source, "to avoid any license contamination" [Maintainer claim, Medium]. Whether the "clean-room" label survives the agent-authorship finding in §4.8 is an open question — see there.

**The window is finite (new in v3).** The whole parity enterprise rests on pandas 2.2.3 staying still: the oracle is pinned, the 1,387 packets assert against it, and the DISCREPANCIES ledger is a diff *against 2.2.3*. pandas 3.0's copy-on-write-by-default and string-dtype changes will move the specification the day it ships, and every release after that is a new differential campaign the harness must fund before the core can claim parity again (trajectory trigger 8). The project is therefore racing a moving specification with a single maintainer and a red conformance gate — the "why now" cuts both ways: the 2.x surface is stable enough to pin *today*, which is exactly why the parity lead is perishable [Inference, Medium].

**Adjacent context:** one entry in the solo-maintainer FrankenSuite program. The asupersync relationship here is the *opposite* of frankenredis's: evaluated and integrated as an optional, off-by-default feature (`fp-runtime/asupersync`, `asupersync 0.5.0`) — and v2 corrects v1's "single example" framing: the gated module is a 1,165-line artifact-transfer subsystem (codec, capability-gated config, Fnv1a integrity verifiers, recovery policies, transport layer with `TransferReport`s), consumed by `fp-conformance` behind the feature gate, plus an `outcome_to_action` mapping (`Outcome` → `DecisionAction`: Ok→Allow, Err→Repair, Cancelled/Panicked→Reject) — corroborated by the module sources and the `fp-conformance` import [Code-verified, High]. The program's thesis appears to be that foundational C/Python infrastructure can be re-derived in Rust with machine-checked parity; frankenpandas is its API-parity flagship.
## 4.2 What it is — repo TL;DR

A 15-crate Cargo workspace (root `Cargo.toml` members enumerated at the pin [Counted, High]) implementing the pandas API in Rust: DataFrame/Series (`fp-frame`, the giant at ~208k claimed lines), a columnar engine with typed `Arc<[f64]>`/`Arc<[i64]>` backings and a bitpacked `ValidityMask` (`fp-columnar`), five typed Index variants plus MultiIndex (`fp-index`), GroupBy with three execution paths (`fp-groupby`), six join directions plus asof (`fp-join`), an expression engine for `eval`/`query` (`fp-expr`), 14+ IO formats (`fp-io`), a differential conformance harness against a pinned live pandas 2.2.3 oracle (`fp-conformance`), a strict/hardened Bayesian runtime policy layer (`fp-runtime`), an ISA-isolated AVX2/FMA dot-product kernel (`fp-dot-kernel`), PyO3 Python bindings (`fp-python`), a TUI scenario harness whose TUI itself is not in-tree (`fp-frankentui`), a benchmark harness (`fp-bench`), and the unified facade (`frankenpandas`). Development velocity is extreme: the contributor graph shows 8,621 commits by the single human maintainer [External, High]; the beads tracker holds 4,051 records, 4,049 closed [Counted, High]. The honest comparison: the zero-`unsafe` invariant is real (all 15 crate roots gated, zero `unsafe` tokens in code [Counted, High]); the "4x faster than pandas" number is a certified-lane geomean with a documented funnel (359 measured lanes → 143 certified) [Code-verified, High].

One-line verdict: **the most methodologically self-scrutinizing codebase assessed in this program to date [Inference, Medium — the in-packet evidence base is the 44,086-line negative-evidence ledger with inline retractions, the certified-lane benchmark gate, 1,387 conformance packets, and the independently re-counted zero-unsafe census; "most" is a comparative judgment over the packets completed so far, not a measured ranking], making API-parity claims its own ledger partly contradicts, under a license that forbids the most likely evaluators from touching it and a contributions policy that refuses outside help.** (NODUS: Explore — see §4.9.)

## 4.3 Repo facts (claim inventory)

Every claim below was verified against the pinned commit on 2026-09-22 (sparse checkout; `artifacts/bench`, `artifacts/phase2c`, and the fuzz corpus were not pulled — claims needing them are marked). Tier flavors as in the header.

| # | Claim | Status | Evidence | Tier, Confidence |
|---|-------|--------|----------|------------------|
| 1 | 15 workspace crates, 648,671 Rust lines / 792 files under `crates/` (567,958 lines / 65 files under `src/`) | demonstrated | `Cargo.toml` members enumerated; `find crates -name '*.rs'` counted. README's dated count ("~587,000 under crates/, ~507,400 under src/, as of 2026-09-03") is 10–12% low — growth since, not contradiction | [Counted, High] |
| 2 | Zero `unsafe`: every crate `#![forbid(unsafe_code)]`, no `unsafe` anywhere | demonstrated | All 15 crate roots carry the gate (4 of them below line 5 — fp-frame's sits under `#![feature(portable_simd)]`); `grep '\bunsafe\b'` over all `.rs` returns only comments/doc-comments | [Counted, High] |
| 3 | 1,387 conformance packet JSON files | demonstrated | `ls crates/fp-conformance/fixtures/packets/ \| wc -l` = 1,387 — exactly the README's number. (README's "1,346" and "1,341" in other spots are stale.) | [Counted, High] |
| 4 | "8700+" tests | demonstrated | 8,954 `#[test]` attributes counted. The FAQ's "5,000+ in-source tests" is conservative-stale, not inflated | [Counted, High] |
| 5 | 143 certified lanes, ~4x geomean vs pandas 2.2.3 | demonstrated (as a generated artifact); the *headline* is partially demonstrated | `artifacts/perf/SCORECARD.md` (generated 2026-09-03 by `gen_perf_scorecard.py`): 143 certified of 359 measured lanes, geomean 3.967x, 10 certified losses listed with thread counts. The gate (live incumbent same invocation, per-arm A/A null, bootstrap median-CI, ELF pinning) is documented in-repo | [Code-verified, High] on the artifact; [Maintainer claim, Medium] on the "exceeds pandas" headline |
| 6 | fp-python: 100% top-level export parity (119/119), 100% core-class coverage (1484/1484) | partially demonstrated | Bindings exist (`fp-python`: PyO3 lib, 2,598-line `.pyi` stubs, differential pytest, maturin wheel builds; "Python wheel & differential tests" CI job green 2026-09-22). But the numerals appear **only** in README prose (5 mentions) — the checked-in drift artifact (`artifacts/api-coverage-drift-2026-05-08.json`, as_of 2026-05-10) attests surface coverage (DataFrame 99.04%, 207/209) but not the headline numerals, and `pandas_api_listing.json` is not in-tree (v4 re-verified: absent at the pin) | [CI-observed, High] on wheel/differential tests; [Maintainer claim, Medium] on the numerals |
| 7 | Strict/Hardened runtime modes, EvidenceLedger, ConformalGuard, RaptorQEnvelope | demonstrated (existence; behavior not executed) | `RuntimePolicy`/`RuntimeMode` (struct, `strict()`/`hardened()` constructors), `EvidenceLedger`, `ConformalGuard`, `RaptorQEnvelope` (manifest + symbol hashes + scrub status + decode proofs) all in `crates/fp-runtime/src/lib.rs` | [Code-verified, High] |
| 8 | RaptorQ-backed artifact durability | demonstrated | Real `raptorq 2.0.1` dependency (with `serde_support`) in `fp-conformance/Cargo.toml`; the envelope format in `fp-runtime` wraps it with manifests and decode proofs. README's "RaptorQ-backed artifact durability" is accurate as stated | [Code-verified, High] |
| 9 | asupersync integration | demonstrated | Optional `asupersync 0.5.0` (`default-features = false`) behind the `asupersync` feature in `fp-runtime`; the gated `pub mod asupersync` is a **1,165-line, 7-file artifact-transfer subsystem** (codec/integrity/recovery/transport: `ArtifactCodec`, `EncodedArtifact`, `Fnv1aVerifier`, `TransferReport`, capability gates, recovery policies) — v1's "single example" undersold it; `outcome_to_action` maps the external `asupersync::Outcome` → `DecisionAction` (Ok→Allow, Err→Repair, Cancelled/Panicked→Reject); consumed by `fp-conformance` (feature-gated import of the codec/integrity/config surface). Off by default — interop, not architecture | [Code-verified, High] |
| 10 | fp-dot-kernel: the one crate compiled with +avx2,+fma | demonstrated (with a documented trap) | Crate docs: per-crate rustflags produced ymm registers with zero `unsafe`; the identical-source +avx2 build once checksummed bit-identical to baseline (both `4957dea0fe3e2ed1`); callers must guard with `is_x86_feature_detected!` or take SIGILL; entry is `#[inline(never)]`/non-generic because inlining silently loses the flags | [Code-verified, High] |
| 11 | 28 numbered divergence entries (16 active, 12 resolved) in DISCREPANCIES.md | demonstrated; one README spot is stale | `grep -c '^### DISC-'` = 28; 16 under "Active", 12 under "Resolved". README's "27 numbered entries" (What's In The Box table) is stale by one. Two entries are substantively stale: DISC-008 still lists "No Python bindings shipped" as active though `fp-python` ships; DISC-007 still says SQL IO is SQLite-only though Postgres/MySQL adapters shipped in 0.3.0 | [Counted, High] |
| 12 | ~4,000 beads, "roughly 80 not closed" (2026-09-02) | stale (in the project's favor) | `.beads/beads.base.jsonl`: 4,051 records — 4,049 closed, 1 in_progress, 1 tombstone. The open backlog shrank from ~80 to ~2, or the base file lags the WAL; either way the README figure is not current | [Counted, High] |
| 13 | 44,086-line NEGATIVE_EVIDENCE.md with inline retractions | demonstrated | `wc -l`; sampled rows include a PERF-CLAIM CORRECTION (a 1.20x "win" re-diagnosed as build variance) and a METHODOLOGY row declaring sub-1.5x ratios UNRESOLVED after measuring 2.6x same-source build variance across CI workers | [Counted + Code-verified, High] |
| 14 | v0.3.0 released (GitHub); 0.2.0 "published" to crates.io 2026-07-27 | partially demonstrated | Ten GitHub Releases (workspace + 9 per-crate, `v0.3.0`, 2026-09-12) [External, High]. crates.io publication **unverified** — crates.io API returns HTTP 403 from this network [Not verified]. No release artifact points at the assessed commit (tags target the earlier `0ad22f91`) | [External, High] / [Not verified] |
| 15 | CI: full matrix green on the pin | **disproven** (for the conformance + all-features legs) | Latest `ci.yml` run (2026-09-22, head `f4d10913`): 19 jobs — 15 green (fmt, tests on 3 OSes, wheel+differential, fuzz-regression, licenses, secret-scan…), 3 red: `conformance` (the live-oracle parity gate), `lint --all-features` (ubuntu + macos); `gates` skipped. `fuzz-nightly` red on the last 5 runs; `fast-gate` green at latest; `wheels.yml` red ×5 at the v0.3.0 tag commit | [CI-observed, High] |
| 16 | Beads-tracked issues, review-mode audit discipline | demonstrated | 4,051 beads [Counted]; AGENTS.md exists as the project's own agent-operating rules (871 lines; includes a named anti-reward-hacking list the maintainer says was "already observed in this suite") [Code-verified, High on existence] | [Counted/Code-verified, High] |
| 17 | Agent co-authorship of the tree | demonstrated (trailer census; line share not measured) | Commits API pages 4–7 (~400 older commits): 398 `Co-Authored-By` hits — Claude Opus 5 (270), Claude (82), Claude Fable 5.1 (32), Opus 5 1M (6), Codex/OpenAI (4), Claude Fable 5 (4). Pages 1–3 (~300 newest commits): 1/300 — the trailer convention lapsed, which is *absence of measurement*, not evidence of human-only authorship | [External, High] on the counts; the clean-room implication is [Inference, Medium] — see §4.8 |
| 18 | "I do not accept outside contributions" | demonstrated | README "About Contributions", verbatim: "I do not accept outside contributions for any of my projects… I'll have Claude or Codex review submissions via `gh` and independently decide whether and how to address them." | [Code-verified (README text), High] |
| 19 | License: MIT + OpenAI/Anthropic rider; rider disclosed in the manifest | demonstrated | LICENSE read verbatim (73 lines; FrankenSuite template + destroy-on-breach clause). Root `Cargo.toml` explicitly declines the SPDX `MIT` label: "not expressible as a standard SPDX expression… no `license = "MIT"` SPDX field is set (it would understate the rider)" — the project discloses the non-OSI status itself | [Code-verified, High] |
| 20 | Five Bedrock Invariants (no unsafe, differential conformance, three-way nulls, zero panics in public APIs, determinism) | partially demonstrated | Invariants 1–3 verified in code/docs (claims 2, 3, three `NullKind` variants in `fp-types` per README FAQ). Invariant 4 (zero panics in public APIs) and 5 (determinism) not executed/verified in this assessment — `cargo doc -D warnings` gate exists in CI per README but was not run here | [Code-verified, High] on 1–3; [Not verified] on 4–5 |

**What the inventory says in aggregate:** claims about *process* (unsafe census, packet counts, test counts, the ledger, the benchmark gate's documented mechanics) verify at High confidence — often exactly (1,387 packets). Claims about *outcomes* ("4x faster", "100% parity", "drop-in replacement") are either funneled through a certification gate the packet must discount (143/359 lanes), asserted only in README prose (119/119), or contradicted by the project's own CI (conformance red) and its own stale ledger entries (DISC-007/008). The README drifts in at least four places (packet counts 1,346/1,341, DISC count 27, bead backlog ~80, LOC 507K) — all small, all in the same direction (stale-low except the parity numerals), and the project's own audit cadence (beads, reality checks) is the mitigation, but the drift rate vs. audit rate is unmeasured [Inference, Medium].

## 4.4 Codebase tour

**Workspace topology (15 members, [Counted, High]):** `frankenpandas` (facade/prelude), `fp-types`, `fp-columnar`, `fp-dot-kernel`, `fp-index`, `fp-frame`, `fp-expr`, `fp-groupby`, `fp-join`, `fp-io`, `fp-conformance`, `fp-bench`, `fp-runtime`, `fp-frankentui`, `fp-python`.

**Dependency posture:** external dependencies were not fully inventoried (no `Cargo.lock` in the sparse checkout — noted in §4.13). Verified individually: `raptorq 2.0.1` (fp-conformance), `asupersync 0.5.0` optional (fp-runtime), `bumpalo` (arena-backed groupby/join per README + `fp-groupby` docs), PyO3 (fp-python), `rusqlite` (default SQL backend). **No tokio** anywhere in the assessed materials — the workspace's no-Tokio policy is load-bearing (ORC APIs fail closed "until a Tokio-free backend lands"; the Postgres adapter is a "pure synchronous wire protocol") [Code-verified, High]. The unsafe surface of dependencies is un-audited here — the zero-`unsafe` invariant covers first-party code only, and the README's own rationale section concedes the point (it discusses resisting `unsafe`-required APIs in `arrow-rs`/`parquet`) [Maintainer claim, High].

**Data flow, end to end:**

- **Types — `fp-types`:** `Scalar`, `DType`, `NullKind` (Null/NaN/NaT as three distinct kinds), `Timestamp`/`Timedelta`/`Period`/`Interval` as proper value types, `SparseDType` scaffolded. Coercion via `common_dtype()`/`cast_scalar()` [Maintainer claim, Medium — per README type-system table; `NullKind` three-variant structure verified in the FAQ].
- **Storage — `fp-columnar`:** `Column` with typed backings (`Arc<[f64]>`, `Arc<[i64]>`, contiguous Utf8 byte buffers — 8 bytes/element for numerics) plus a lazily materialized `Vec<Scalar>` (16 bytes/element + tag) built only when a scalar-per-element API is called; `ValidityMask` as a bitpacked `Vec<u64>` (1 bit/element, vs pandas' 8-byte nullable dtype) enabling word-level null algebra [Maintainer claim, Medium]. AG-10 typed-array vectorization dispatches binary ops to `&[f64]`/`&[i64]` slices for LLVM auto-vectorization — SIMD without `unsafe`, enabled by the workspace's `portable_simd` nightly feature [Code-verified, High — `fp-frame` root carries `#![feature(portable_simd)]`].
- **Alignment — `fp-index` → AACE:** every binary op routes through an explicit `AlignmentPlan` (`union_index, left_positions, right_positions`) before materialization; the plan is recorded in the `EvidenceLedger` with Bayesian posterior and expected losses. A fork-wide sweep retrofitted 70+ callsites to propagate index/axis names [Maintainer claim, Medium]. AG-05 leapfrog triejoin does N-way index union in one O(n log n) sorted-merge pass [Maintainer claim, Medium].
- **Frame — `fp-frame` (~208k claimed lines, the giant):** DataFrame/Series, 960+/800+ claimed methods, accessors (str 50+, dt 25+), windows (rolling/expanding/ewm/resample), reshape, MultiIndex, a native deterministic SVG/HTML plotting renderer (11 plot kinds, zero-dependency pure safe Rust) [Maintainer claim, Medium — method counts not independently re-counted; the plot renderer's `plot_render.rs` header confirms the no-dependency/no-unsafe claim, Code-verified, High].
- **GroupBy — `fp-groupby`:** three execution paths (dense Int64 for range ≤ 65,536, Bumpalo arena, HashMap fallback) with property tests asserting bitwise equivalence [Maintainer claim, Medium].

**Worked trace — `df.groupby('k').sum()` (new in v2, code-verified).** `fp-frame` hands the keys and values columns to `fp-groupby`, which dispatches before computing anything. First probe: `try_groupby_sum_dense_int64_slices` checks whether keys and values are all-valid Int64 buffers whose key span fits `DENSE_INT_KEY_RANGE_LIMIT` (`const DENSE_INT_KEY_RANGE_LIMIT: i128 = 65_536`, line 1000); if so, the dense path accumulates i128 totals with a first-seen tape and ascending bucket emission *directly over the raw `&[i64]` slices* — explicitly "avoiding full-column Scalar materialization," i.e., dodging the 16-byte-per-element lazy `Vec<Scalar>` tax from `fp-columnar` [Code-verified, High]. If the span check fails the function returns `None` and the generic path takes over: float keys are normalized with `to_bits()` (NaN canonicalized to a single bit pattern, line 653) so key equality is bitwise-deterministic; intermediates allocate in a `bumpalo::Bump` arena when `use_arena: true` within `arena_budget_bytes`; the result struct records `used_arena` so callers can observe which path ran [Code-verified, High]. The three paths are pinned together by dispatch tests (`dense_median_dispatch_matches_generic_bits`; AG-08-T: "Int64 keys with span > 65536 forces generic path") asserting the dense and generic paths produce identical results [Code-verified, High]. This is the packet's "architecture reconstructed from code" in miniature: the zero-`unsafe` constraint forces the performance story into *representation choices* (typed slices, bit-normalized keys, arenas) rather than pointer tricks, and the dispatch boundary is where the benchmark-relevant performance actually lives — which is also why groupby is the category where the certified census drops the most lanes on variance (38/90 DROPPED_HIGH_CV, §4.5): the three-path dispatch multiplies the measurement surface [Inference, Medium].
- **Joins — `fp-join`:** six directions + asof with `tolerance`/`by`/`allow_exact_matches`, `validate` modes, indicator/suffixes [Maintainer claim, Medium].
- **Expressions — `fp-expr`:** `eval`/`query` with correct precedence, `@local` bindings, backtick columns, chained-comparison pairwise AND [Maintainer claim, Medium].
- **IO — `fp-io`:** 14+ formats with deep option matrices (CSV's option list alone spans ~20 knobs); SQL via a generic `SqlConnection` trait (rusqlite default, MySQL behind `sql-mysql`, Postgres via a Tokio-free synchronous wire protocol behind `sql-postgresql`); ORC fail-closed under the no-Tokio policy; Pickle round-trips inside a FrankenPandas envelope (not interchangeable with pandas pickles — documented) [Maintainer claim, Medium; the ORC fail-closed and pickle-envelope caveats are Code-verified in the README table, High].
- **Runtime — `fp-runtime`:** `RuntimePolicy` (Strict fails fast; Hardened runs the Bayesian Allow/Reject/Repair decision under an asymmetric loss matrix — false-allow priced 200x false-reject-when-incompatible), `EvidenceLedger` (append-only JSONL with priors, posteriors, Bayes factors, evidence terms), `ConformalGuard` (split-conformal distribution-shift detection on the conformance gate's fixtures), `RaptorQEnvelope` (FEC-protected durable artifacts with scrub reports and decode proofs) [Code-verified, High on existence].
- **Conformance — `fp-conformance`:** 1,387 packet JSONs [Counted, High]; the live-oracle batch (pinned pandas 2.2.3 in `.venv-oracle`) runs in CI's daily batch; DISCREPANCIES.md's 28 entries; the `raptorq` dependency that actually encodes the repair symbols [Code-verified, High].
- **Bench — `fp-bench` + `benches/vs_pandas_harness.py`:** the harness behind the certified-lane numbers; `artifacts/bench/` holds 960 JSON rows at the pin [Counted, High], 31 MB — the raw corpus the scorecard certifies over.
- **Python — `fp-python`:** PyO3 bindings, 2,598-line `.pyi` stubs, differential pytest over the packet fixtures, maturin wheel builds [Code-verified, High on existence].
- **TUI — `fp-frankentui`:** dashboard snapshot value types + E2E scenario harness; "the TUI itself is not in-tree" [Maintainer claim, High].

**Concurrency model:** no rayon, no tokio. Hot paths fan out with `std::thread::scope` at 147 occurrences across 8 files, sized from a cached `available_parallelism()`; string kernels use one persistent worker pool [Maintainer claim, High — README L2641]. The structural cost is the per-call spawn (~350–420µs on the reference host), which is the maintainer's causal account of the 100k-row losses — and the tracked fix is a shared pool [Maintainer claim, High].
## 4.5 Benchmark and conformance audit

**Maintainer-produced numbers.** The instrument is the certified census over `artifacts/bench/` (960 JSON rows at the pin [Counted, High], 31 MB). The generated scorecard (2026-09-03) reports: 359 lanes, **143 certified**, overall certified geomean **3.967x** (pandas time ÷ FrankenPandas time), **10 certified losses** [Code-verified, High]. The certification gate, documented in the scorecard header and `scripts/current_loss_census.py`: the incumbent must run **live in the same invocation**, a per-arm **A/A null control**, a **balanced ABBAABBA square**, and a **bootstrap median-CI gate** whose clauses must all pass; rows failing the gate (NULL_UNDECIDABLE, DROPPED_HIGH_CV, CONTRACT_INVALID) are **counted, not averaged**; one row per lane (newest certified); each row carries an **ELF SHA-256 pin**, and surviving losses whose ELF isn't the lane's newest are flagged STALE ELF ("re-measure before acting") [Code-verified, High]. The 10 certified losses are listed with both arms' thread counts — e.g. `df_transpose_full_materialize` at 0.004x (a 2-D block-storage floor) and `drop_duplicates` at 0.840x *with FrankenPandas on 10 threads against pandas on 1* — losing with more threads, reported anyway [Code-verified, High]. The five named optimization rounds are explicitly fenced: "Every figure in the table below is a FrankenPandas-vs-FrankenPandas improvement… **None of these are comparisons against pandas**, and none should be read as one" [Code-verified (README), High] — the anti-misrepresentation note frankenredis's packet had to wish for.

**The maintainer's own limits on these numbers.** `docs/NEGATIVE_EVIDENCE.md` row 118 (METHODOLOGY): repeated rebuilds of *identical source* across the project's remote build workers swung wall-time **2.6x** (heterogeneous `target-cpu`), so "any A/B whose two binaries were built on DIFFERENT workers is invalid" and **sub-1.5x ratios are UNRESOLVED, not wins** — a standing, self-imposed discount on marginal claims [Code-verified, High]. Row 112 (PERF-CLAIM CORRECTION): a claimed 1.20x where/mask win was re-diagnosed in the same row as build variance and reverted ("trust sub-1.5× 'wins' only with a consistent same-pin A/B") [Code-verified, High]. The census script's docstring documents two past self-foolings (worst-ever vs latest row selection; measuring a dead ELF) and the discipline that replaced them [Code-verified, High].

**The funnel, aggregated (new in v2).** The scorecard's per-category "Latest verdicts (all lanes)" lets the certification funnel be counted, not gestured at [Code-verified, High]. Across all 359 lanes: **FASTER 201, SLOWER 20, NULL_UNDECIDABLE 74, DROPPED_HIGH_CV 62, PARITY 2.** Of the 143 certified lanes, 133 are certified wins and 10 are certified losses; the 216 uncertified lanes decompose as 68 faster-but-not-certified, 10 slower-but-not-certified, 74 undecidable, 62 high-CV, 2 parity. Two readings. The charitable one: the gate is *strict* — 68 lanes currently read FASTER but are withheld from the certified set because they failed a clause of the median-CI gate, which is the apparatus working as designed (withholding wins, not manufacturing them). The uncharitable one: the 3.967x geomean is computed over the 133 certified wins plus 10 certified losses — i.e., over the decidable subset — while 136 lanes (74 + 62) were excluded for undecidability or variance rather than averaged, and another 78 faster/slower lanes sit outside certification. The geomean is therefore a *decidable-subset* statistic, and its selection pressure runs in exactly one direction: toward lanes where the measurement apparatus could reach a verdict — a *decidability* filter, not a wins-only filter (whether that filter flatters the number is tested, not assumed, in the mechanism check below). The worst category is groupby — 38 of 90 lanes DROPPED_HIGH_CV — which is also the category where the maintainer's arena-optimization story lives; the optimization narrative and the measurement failure share a category [Code-verified, High on the counts; the two readings are Inference, Medium]. Also new: the scorecard states the oldest certified row still standing is 2026-08-15 and the newest certified rows are 2026-09-01 — the tree moved for three weeks on top of certified numbers, and the STALE ELF flags mark exactly this staleness [Code-verified, High].

**Does the funnel flatter? A mechanism check (new in v3).** The uncharitable reading above needs a mechanism — *why* would decidability correlate with FrankenPandas wins? The per-category data lets us test the obvious candidate (the gate drops exactly the lanes where FP loses). It doesn't hold cleanly: the two categories with certified losses — math_unary (6 certified losses) and dataframe_ops (4) — have the *lowest* drop rates (0/65 and 10/83), i.e., losses certify just fine where variance is low; the gate is not hiding them. But groupby cuts the other way and deserves the attention: 90 lanes → 9 certified, 10 SLOWER-but-uncertified, 38 DROPPED_HIGH_CV — in the category of the project's flagship arena optimization, the certified 4.718x geomean rests on 9 lanes while 10 lanes reading slower sit outside certification and 38 never got a verdict. And the 68 faster-but-uncertified lanes across the corpus are the mirror image: wins the gate withheld because a median-CI clause failed. The honest position: the gate's conservatism is *symmetric* — it withholds wins and losses alike — which is the best defense of the funnel's integrity, but it does not rescue the geomean as a population statistic. A 3.967x over the decidable subset, where the noisiest category certifies 10% of its lanes, is a statement about the measurable, not the typical [Code-verified, High on the counts; Inference, Medium on the conclusion].

**What would survive an independent rerun.** The *direction* of the large structural wins (the O(n²)→O(n) sweep, the AXPY reorder, arena-backed groupby) is the most defensible — large deltas survive build variance by the maintainer's own rule [Inference, Medium]. The 3.967x geomean would likely *not* reproduce as a number: it is a decidable-subset statistic (133 certified wins + 10 certified losses of 359 lanes; 216 lanes uncertified), the newest certified row per lane dates to 2026-09-01 while the tree moved on (STALE ELF flags say so), and cross-worker build variance is un-controlled in the corpus [Inference, Medium]. Which numbers are load-bearing for the thesis? The parity numbers (1,387 packets, the live-oracle batch), not the speed numbers — the project's own "semantic parity over speed" invariant says so, and speed is where the maintainer applies the heaviest discount to himself [Inference, Medium].

**Independent numbers:** none known. A web search for independent benchmarks, reviews, or production users of frankenpandas returned nothing beyond the repository on 2026-09-22 [External, High within recall caveats]. No independent rerun of the certified census is known to exist.

**Reproduction cost:** the bench corpus alone is 31 MB in-tree; a rerun needs the pinned nightly toolchain (`rust-toolchain.toml`), a pandas 2.2.3 oracle venv, and — per the maintainer's own methodology finding — a *single pinned host* (or `RUSTFLAGS=-Ctarget-cpu=x86-64-v3` pinning) to keep build variance from swamping marginal lanes. The full corpus at the pin is 960 rows across 359 lanes; the scorecard's "latest verdicts" show the gate is the expensive part (e.g. groupby: 38 of 90 lanes DROPPED_HIGH_CV) [Counted/Code-verified, Medium].

## 4.6 Comparison: who owns the lane

**pandas 2.x (the incumbent)** owns the lane outright: the API being reimplemented, the oracle being tested against, ~15 years of edge-case accumulation, and the entire Python data ecosystem. FrankenPandas's pitch concedes this implicitly — every claim is *relative to pandas* (parity packets, vs-pandas lanes), which makes pandas both the specification and the ceiling: any pandas release can move the parity surface (the oracle is pinned at 2.2.3; pandas 3.0's copy-on-write and string dtype changes would be a new differential campaign) [Inference, Medium].

**Polars (Rust, own API)** owns the "fast DataFrame in Rust" lane in buyer perception: lazy execution, a real query planner, streaming/out-of-core, rayon parallelism — everything frankenpandas's honest-comparison section admits it lacks ("Polars optimizes across chains; FrankenPandas doesn't"; "production-grade parallelism today (Polars uses rayon by default; FrankenPandas fans out with std threads per call and has no shared pool yet)") [Maintainer claim, High]. Polars wins on performance engineering; it does not offer pandas API parity, which is the whole point of differentiation [Inference, Medium].

**DuckDB (SQL, columnar)** owns the "fast analytics without leaving your language" lane via SQL, with Arrow-native zero-copy interop and a mature optimizer. The README's honest version positions them as composable, not competitive: "use DuckDB / Polars for the heavy filter/aggregate pass, hand the result to FrankenPandas (via Feather / Parquet) for the pandas-shaped transformations downstream" [Maintainer claim, High]. That composition story is also a concession: frankenpandas is the *downstream* tool in its own architecture diagram [Inference, Medium].

**The unoccupied lane:** a memory-safe, pandas-API-complete DataFrame library in Rust with machine-checked parity evidence and an auditable alignment/runtime policy layer has no direct occupant — Polars/DuckDB chase speed with different APIs. The "pandas in another language" attempts stay inside Python and each fails the lane on a different axis: **Modin** (distributed pandas API over Ray/Dask) offers API familiarity but inherits Python's memory/GIL costs and has no auditability story; **cuDF** (RAPIDS) is a GPU-bound pandas API — a hardware axis, not a safety or auditability axis, and it ties the buyer to NVIDIA; **Vaex** (out-of-core lazy DataFrames) competes on scale-per-machine with its own API dialect, not on parity-plus-proof [Inference, Medium — positioning from public project descriptions, not re-verified in this round]. None of them sells *evidence*: machine-checked parity packets, a published negative-evidence ledger, a fail-closed Bayesian runtime policy. Whether any buyer pays for API-parity-plus-auditability over Polars' speed-plus-new-API is the unproven bet the whole project rests on.

**Who would pay (demand-side, new in v2).** Three candidate buyers, in descending order of plausibility [all Inference, Low — no demand signal was found, §4.13]: (1) **Regulated-industry data pipelines** (finance, pharma, aerospace-adjacent) where a pandas-shaped transformation layer with an EvidenceLedger of every alignment/coercion decision is auditable in a way `import pandas` is not — the buyer here pays for the *ledger*, not the speed; (2) **Agent frameworks executing untrusted pandas snippets**, where fail-closed modes (Strict rejects, Hardened logs-and-recovers) bound what a runaway agent can silently corrupt — but this buyer is currently *excluded by name* in the rider, which is the packet's central irony; (3) **Rust shops** that want pandas semantics without a Python interpreter — the smallest pool, since they could also just learn Polars. The honest read: the lane is unoccupied because it may be unoccupiable at commercial scale — auditability is a feature buyers praise and decline to pay for — and the rider plus the no-contributions policy exclude the two buyer pools (AI labs, agent frameworks) most likely to value it [Inference, Medium].

## 4.7 Skeptic's take

*Citation convention: weaknesses are numbered 1–8 below; elsewhere "§4.7.N" means weakness N.*

1. **[FATAL] The rider poisons the well it drinks from.** Barring OpenAI/Anthropic *and their affiliates and anyone acting for them* from use, benchmarking, analysis, indexing, and training-data incorporation — with automatic termination plus a *destroy-all-copies* clause harsher than frankenredis's — doesn't just block two companies. The "acting directly or indirectly for … or for the benefit of" prong creates legal uncertainty for any contributor near the AI supply chain, and a project whose moat is *evidence* forbidding evidence-gathering is self-negating. The rider's "analyzing" bar arguably covers this very assessment [Inference, High — legal conclusion, not legal advice].
2. **[FATAL] The contributions policy is a second, independent ceiling.** The README's "About Contributions" refuses outside contributions outright ("I do not accept outside contributions for any of my projects… I'll have Claude or Codex review submissions via `gh`"). Even if the rider were lifted tomorrow, there is no on-ramp for a second committer — the bus factor is 1 by *policy*, not just by circumstance [Inference, High]. No assessed FrankenSuite repo states the closed-door policy this explicitly.
3. **[HIGH] The headline number is a certified-lane geomean, and the certification funnel is doing work.** 359 measured lanes → 143 certified; the gate drops high-CV and undecidable lanes rather than averaging them (honest, but selective); the maintainer's own methodology ledger declares sub-1.5x ratios UNRESOLVED after measuring 2.6x same-source build variance. A 3.967x geomean over the decidable subset, with stale ELFs in the corpus and no independent rerun, is a process artifact as much as a result — full funnel aggregation and the flattery-mechanism check in §4.5 [Inference, Medium].
4. **[HIGH] The parity gate is red and the ledger drifts.** The `conformance` CI job (the live-oracle parity gate) failed on the latest run; `--all-features` lint is red on two legs; `fuzz-nightly` is red five runs straight [CI-observed, High]. Meanwhile DISCREPANCIES.md — the divergence ledger that is supposed to be the source of truth — carries two substantively stale entries (DISC-007's "SQLite-only" after the Postgres/MySQL adapters shipped; DISC-008's "no Python bindings" after `fp-python` shipped) [Code-verified, High]. A drifted ledger in a project whose pitch is *auditability* is a structural liability, and the README drifts in four more places (claim inventory, §4.3).
5. **[MEDIUM] The Python-parity numerals are README-only.** 119/119 and 1484/1484 appear five times in README prose and nowhere else — the only checked-in instrument is a drift artifact filename-dated 2026-05-08 (`as_of` 2026-05-10) that attests DataFrame surface coverage (99.04%), not the headline numerals. The bindings demonstrably exist and wheel CI is green, but the exact numerals that carry the "100% drop-in" claim are unattested at the pin [Inference, Medium].
6. **[MEDIUM] The concurrency story is "spawn a thread per call and eat 400µs".** 147 `thread::scope` fan-out sites, no shared pool (the fix is tracked, not landed), no rayon, no query planner, no out-of-core. The README's own honest comparison concedes Polars wins on "production-grade parallelism today." The 100k-row losses are the visible symptom; the architecture is single-threaded-with-sprinkles [Inference, Medium].
7. **[HIGH] Agent co-authorship murkies the "clean-room" label.** ~398 agent-co-authored commits in the sampled older window (claim 17) — 394 of them Claude-family from Anthropic, a named Restricted Party; 4 from OpenAI's Codex. Promoted from MEDIUM in v2: the clean-room posture is load-bearing for the FrankenSuite's program-level framing (all 44 repos are "clean-room reimplementations"), and a tree materially co-authored by models plausibly trained on pandas puts that framing — not just this repo's marketing — under the output-similarity test. If models plausibly trained on pandas wrote material parts of the tree, the clean-room question is copyright's output-similarity test, which the rider's no-training clause cannot answer — and the trailer convention *lapsed* for the newest ~300 commits, so the recent AI share is unmeasurable [Inference, Medium — full analysis in §4.8].
8. **[MEDIUM] The embeddable story is vendored, not versioned.** Every adoption path in the docs is a relative `path` dependency; the workspace pins a nightly toolchain and 2024-edition features (`portable_simd`); and adopters inherit the rider. There is no registry-artifact adoption story verified in this assessment (crates.io unverifiable from here) [Inference, Medium].

**Bear-case steelman (strongest counter-case):** FrankenPandas is a single-maintainer methodology performance about *how* to re-derive an API, not a pandas replacement. The 4x number is a certified-subset geomean its own author discounts below 1.5x; the parity gate is red; the two ledgers that constitute its auditability story both drift; the license forbids the only entities with the resources to validate a 649k-line safety claim from even analyzing it; the contributions policy refuses the second committer who could fix any of this; Polars already owns "fast DataFrames in Rust" with real users and a query planner; and pandas 3.0 will move the parity surface the day it ships. The rational market response is to mine it for methodology (certified-lane benchmarking, negative-evidence ledgers, differential-against-oracle conformance, fail-closed runtime policy) and never depend on it. Its most likely end state is abandonment at the maintainer's next context switch — leaving 649k lines of unauditable-by-license Rust whose best-equipped evaluators may not benchmark, analyze, or train on it. [Inference, Medium — deliberately uncharitable; the counter-evidence is the repair velocity — 4,049 closed beads — and the ledger culture that documents its own retractions.]
## 4.8 Maintenance & succession

**Bus factor: 1 — by policy, not just circumstance.** One human maintainer (Jeffrey Emanuel); the contributor graph shows 8,621 commits by him plus 10 by dependabot[bot] and nothing else [External, High]. The README's "About Contributions" makes the bus factor a *decision*: "I do not accept outside contributions for any of my projects. I simply don't have the mental bandwidth to review anything, and it's my name on the thing… I'll have Claude or Codex review submissions via `gh` and independently decide whether and how to address them" [Code-verified (README text), High]. No succession plan, no second committer, no foundation. The closest thing to institutional knowledge transfer is the in-repo bead tracker (4,051 records, 4,049 closed [Counted, High]) plus the evidence ledgers — process artifacts that outlive a maintainer, and the only onboarding story a successor would inherit. Review depth is unassessed: whether any of the 4,049 closed beads saw a second pair of eyes, and what "review-mode audit" meant in April 2026, were not examined [Not verified].

**Agent authorship vs the "clean-room" posture.** The tree is materially agent-written in its older bulk: in commits API pages 4–7 (~400 commits), 398 carry `Co-Authored-By` trailers — 394 Claude-family (Opus 5, Claude, Fable 5.1, Opus 5 1M-context) and 4 Codex/OpenAI [External, High]. No agent is ever the primary author in the sampled windows [External, High] — but co-authorship at that density, in the commits that built the core, means agents are not marginal contributors. Three consequences: **(1)** the "clean-room" label is legally murky in the same way as frankenredis's — *if* models plausibly trained on pandas wrote material parts of the tree, the question that matters is copyright's output-similarity test, not tool purity; the rider is a licensee-facing term that cannot bind its own author, so its no-training clause is legally inert as to the author's own conduct — the irony (barring as licensee conduct what was practiced as authorship) and the epistemic point (§4.10.8's analog) stand [all Inference, Medium — legal analysis, not legal advice]; **(2)** the rider's "analyzing" bar arguably covers this very assessment — same recursion as the redis packet, with the same rebuttal (nothing establishes the analyst acts for a Restricted Party) [Inference, Medium]; **(3)** quantification limits are worse here than in the redis packet: the trailer convention *lapsed* — 1/300 in the newest three pages — so trailer counts can't even bound the recent AI share; full-tree line-level attribution via `git blame` was not performed and is the open question in §4.13.

**Steelmanning the rider (new in v2 — the maintainer's own policy record).** The packet's v1 framing treated the rider as an inexplicable self-sabotage. That was unfair: the maintainer has *considered* the license question explicitly. Bead `frankenpandas-dio8` ("license: custom MIT+Rider is non-standard for Rust ecosystem — dual MIT/Apache-2.0 is convention (policy decision)", status: **closed**) laid out three options — (1) keep as-is, explicitly framed as "if the rider is intentional policy against certain AI providers' training-data use, keep it and accept the corporate-review friction"; (2) dual-license MIT OR Apache-2.0, dropping the rider; (3) a triple-license variant — and the bead was closed with the rider intact [Code-verified (bead text), High]. Two readings, stated explicitly (v3 correction — v2's "option 1 won" overread the record): **(a) deliberate retention** — the maintainer reviewed the three options and kept the rider as intentional policy against AI training-data use; **(b) retention by default** — the bead is titled [LOW], its severity section calibrates it as LOW "because the question needs a human policy decision," and its `close_reason` explicitly *defers* rather than records a choice ("reopening / re-filing is the right path if the maintainer wants to record a choice in repo rather than here") [Code-verified (bead record), High]. The packet commits to (a)-leaning at [Inference, Medium]: the bead was closed 2026-04-23 and the rider still stands in LICENSE at the 2026-09-22 pin — five months of post-review retention is evidence of tolerance, not of decision, but tolerance sustained that long past explicit review is functionally close to policy. What is *not* claimed: that a deliberate choice was ever recorded — it wasn't. That reframes the §4.7.1 "sabotage" charge: the rider serves its stated defensive intent (deny 649k lines of curated systems code to rival training corpora — and note the §4.7.7 finding: the older bulk of the tree is heavily Claude-co-authored, so barring Anthropic from training on code Claude helped write is at least a *coherent* motive, whether or not it is legally sound) [Inference, Medium]. The sabotage verdict survives anyway, on narrower grounds: whatever it defends against, the rider taxes exactly the currency the project's mission runs on — *evaluation*. A project whose moat is evidence (parity packets, certified lanes, published retractions) forbidding benchmarking, testing, analyzing, and indexing by the entities best equipped to do it is self-negating regardless of motive; and the "acting directly or indirectly for … or for the benefit of" prong creates legal uncertainty for contributors nowhere near OpenAI or Anthropic [Inference, Medium — legal analysis, not legal advice]. **Durable ceiling, not merely blocking (v3: certainty softened one notch):** the rider was adopted at inception (2026-02-18), survived explicit review without being dropped, and only the sole copyright holder can remove it — v3 treats it as a *durable-by-choice* ceiling: functionally constitutive, though the "by choice" is inferred from five months of post-review retention rather than a recorded decision [Inference, Medium]. It can still be removed (one person's decision), but there is no evidence the decision is forthcoming, and the packet's ring logic no longer assumes it (see §4.9, §4.12).

**Maintenance load signals:** extreme velocity (8,621 commits since 2026-02-13 [External, High]; ~4,000 beads closed since the April review-mode audit [Counted, High]) cuts both ways — dedication and a workload no successor inherits casually. The daily CI batch is the canary: the fast-gate exists precisely because the full matrix "produced 0 green runs in 1,200 attempts" and left main with no enforced compile check for days (the `fp-types` test module sat uncompilable 2026-08-30 → 2026-09-02 "with nobody noticing") [Maintainer claim, High — the fast-gate workflow's own header comment]. That incident is the maintainer's own evidence that velocity without gates rots.

**License — the rider, scoped precisely (LICENSE read verbatim at pin, 73 lines; adopted 2026-02-18, five days after repo creation — CHANGELOG [Code-verified, High]):**

- **Title:** "MIT License (with OpenAI/Anthropic Rider)" — and the root `Cargo.toml` explicitly refuses the SPDX `MIT` label for it: "The LICENSE file is MIT plus a non-MIT rider on AI-assistant usage, which is not expressible as a standard SPDX expression… no `license = "MIT"` SPDX field is set (it would understate the rider)" [Code-verified, High]. The project classifies its own license as non-SPDX-expressibly-MIT.
- **Restricted parties:** "OpenAI, L.L.C.; Anthropic, PBC; any of their respective Affiliates; and any person or entity acting directly or indirectly on behalf of, for the benefit of, or under the direction of any of the foregoing (including any officer, director, employee, contractor, agent, consultant, service provider, or representative)." (verbatim)
- **Grant:** "no rights are granted to any Restricted Party… null and void absent the express prior written permission of Jeffrey Emanuel." (verbatim, condensed)
- **Scope:** bars providing or permitting access to "the Software or any derivative work" to or for a Restricted Party.
- **"Use" is defined expansively** — "copying, modifying, merging, publishing, distributing, sublicensing, selling, transferring, making available, hosting, deploying, executing, benchmarking, testing, analyzing, indexing, or incorporating the Software or any Derivative Works into any dataset, training corpus, evaluation harness, or pipeline for machine learning or other automated systems." (verbatim)
- **Remedy:** breach "automatically and immediately terminates" the permissions; "you must immediately cease all use and distribution… and **destroy all copies under your control**" (verbatim — harsher than frankenredis's rider, which lacks the destruction clause); injunctive/equitable relief plus attorneys' fees reserved; the rider is a "precondition to exercising any rights" and must be passed along "unmodified" on distribution.
- **Classification:** non-OSI, source-available. Named-party discrimination fails OSD §5; the benchmarking/training/analysis bar fails OSD §6 [Inference, High — legal conclusion, not legal advice].

**Succession consequence:** even a willing acquirer inherits a license that shrinks the contributor and evaluator pool *and* a contributions policy that refuses contributors; removing either requires the sole copyright holder's action — the same single point of failure as everything else [Inference, Medium].

## 4.9 NODUS factsheet

| Criterion | Score | One-line justification |
|---|---|---|
| Technology readiness | **TRL 4–5** | Lab-validated components (1,387 differential packets vs a pinned live oracle, 143-lane certified benchmark gate, 8,954 tests, maturin wheels building in CI) with zero production exposure, a red parity gate at the latest CI run, README-only parity numerals, and no independent validation — the range, not a point, is honest [Inference, Medium] |
| Strategic relevance | 3/5 | A memory-safe pandas-API core is central to the FrankenSuite infrastructure thesis, but the rider blocks the program's own engagement with the software and the no-contributions policy blocks institutional capture [Inference, Medium] |
| Impact potential | 2/5 | The pandas-API-in-safe-Rust lane is genuinely unoccupied, but no buyer signal exists, the incumbent (pandas) owns the specification, and Polars owns buyer perception of "fast DataFrames" [Inference, Medium] |
| Implementation feasibility | 2/5 | 649k lines, one maintainer *by policy*, nightly toolchain, no release artifact at the pin — any adopter inherits the entire tree plus the rider [Inference, Medium] |
| Time to mainstream | 2/5 | Years at best: rider removal, a contributions on-ramp, a tagged release at a green pin, and independent validation must all happen first, in that order [Inference, Low] |
| Collaboration potential | 1/5 | The rider bars the likeliest collaborators (AI labs and their agents) from even analyzing the code; the contributions policy refuses everyone else; bus factor 1 [Inference, Medium] |

**Ring: Explore** [Inference, Medium]. The Rulebook's default for substantive-but-unproven; the factsheet above is the scaffolding the ring stands on. v2/v3 revise the v1 "advancement blocker" framing: the rider is not a blocker awaiting removal — adopted at inception (2026-02-18), it survived explicit review (bead `frankenpandas-dio8`, closed 2026-04-23) without being dropped, and only the sole copyright holder can remove it (§4.8 — deliberate-retention vs retention-by-default adjudicated there at [Inference, Medium]). It is therefore priced as a **durable ceiling**: recorded in §4.8, priced into feasibility (2) and collaboration (1), and wired to trigger 3 — but the ring logic no longer assumes trigger 3 will fire. The contributions policy is the second ceiling (§4.7.2, trigger 4): **Explore-with-a-double-ceiling**, currently un-advanceable past Explore while both stand, with no evidence either will lift. No Rulebook amendment; v1.0 stands.

**Methodology fit (for the FrankenSuite):** adopt the patterns, not the package — the certified-lane benchmark gate, negative-evidence ledgers with inline retractions, differential-against-oracle conformance packets, bead-tracked issue discipline, and the fail-closed Bayesian runtime policy are directly importable into how the program evaluates the other 42 repositories. Do not depend on the software: the rider, the red parity gate, the unattested parity numerals, and the no-contributions policy disqualify it as a dependency, benchmark target, or training-data source. If the program ever needs a DataFrame engine, Polars dominates on every adoption criterion except pandas-API fidelity — and the rider disqualifies frankenpandas regardless [Inference, High].

---

## 4.10 Wardley placement

Placing the *components*, not the repo:

- **Commodity — the pandas API surface and the CSV/Parquet/Arrow formats.** Inherited standards; API-parity is table stakes, not differentiation — and the specification is owned by the incumbent, which can move it (pandas 3.0) [Inference, Medium].
- **Custom-built, approaching early Product — the 649k-line pandas core re-derivation** (`fp-frame`, `fp-columnar`, `fp-index`, the IO matrix). Custom-built on process rigor alone; not yet product (no release artifact at the pin, red parity gate, pre-1.0 API instability admitted). Moves right with a tagged release at a green pin plus an independent benchmark; decommissions (becomes a monument) if Polars ships a pandas-compat layer or pandas 3.0 moves the parity surface faster than the differential harness absorbs it [Inference, Medium].
- **Custom-built, closest to Product — the differential-against-oracle conformance harness + certified-lane benchmark machinery.** 1,387 packets, the live-oracle batch, SHA-pinned golden fixtures, the bootstrap-CI certification gate with stale-ELF re-measure flags. The most *executed* component in the tree — used operationally in-repo. Moves right if extracted as a standalone conformance/benchmark framework other projects can adopt; stays repo-local tooling, invisible outside the repo, while the rider blocks the likeliest adopters [Inference, Medium].
- **Genesis — AACE + the Bayesian runtime policy (Strict/Hardened + EvidenceLedger + ConformalGuard).** No incumbent ships "preserve index-alignment semantics exactly *and* log every ambiguous decision with a posterior" as a first-class mode switch — checked 2026-09-22: the closest Polars ecosystem artifact is `polars-lineage`, a third-party PyPI PoC for column-level lineage extraction, which tracks *where data came from*, not *why an ambiguous decision was made* — a different problem [External, Medium]. Moves right if the ledger becomes a reviewed, origin-bound provenance mechanism — or stagnates as a single-maintainer log format if the maintainer context-switches [Inference, Medium]. What this Wardley says that §4.11 doesn't: the *components* have divergent fates — the conformance/benchmark machinery is the most extractable (it could outlive the repo as a standalone framework), while the pandas core is the least (649k lines, one maintainer, no release) — so "the project's" trajectory is really four trajectories, and the methodology-export one is the only one with a rightward path that doesn't require the rider to lift.

**Residual product gaps (what's missing / risky):**

1. **No lazy evaluation / query planner** — explicitly "Medium" on the roadmap; Polars' core advantage is uncontested [Maintainer claim, High].
2. **No shared thread pool** — the tracked fix for the 100k-row losses; 147 fan-out sites still pay ~350–420µs per call [Maintainer claim, High].
3. **ORC fail-closed, HDF5 snapshot-only, GBQ/SPSS deferred** — the IO long tail [Maintainer claim, High].
4. **Parity-tail DISC entries open** — DISC-011 (nullable Int64, the 97-fixture root cause), DISC-006 (MultiIndex advanced ops), DISC-009 (sparse storage still dense) [Maintainer claim, Medium].
5. **The fp-python numerals unattested** — §4.7.5.
6. **README + DISCREPANCIES drift as a standing risk** — §4.7.4; the beads audit cadence is the mitigation, drift-rate-vs-audit-rate unmeasured [Inference, Medium].
7. **The 60-day abandonment tripwire** (see trigger 9) — decay is the default for bus-factor-1-by-policy infrastructure [Inference, Medium].
8. **pandas 3.0 moves the specification** — every new pandas release is a new differential campaign the harness must fund before the core can claim parity again [Inference, Medium].

## 4.11 Trajectory (12 / 24 / 60 months)

All horizons **[Inference]** — forecasts, not findings; confidence Low unless noted.

**12-month base case [Inference, Medium]:** velocity continues; the parity tail shortens bead by bead; the shared thread pool lands and the 100k-row losses get re-measured (next-step 5's done-when); the conformance gate flips green intermittently but the `--all-features` lint legs stay red; no contributions on-ramp appears; the rider stands. The methodology gets mined (certified-lane gates, negative-evidence ledgers appear in sibling projects); the software gets no production users. Ring: Explore, unchanged.

**24-month base case [Inference, Low]:** one of two shapes. Either the program compounds — a tagged release at a green pin, the pool fix validated, the DISCREPANCIES ledger de-drifted — and the packet's checklist starts scoring above zero; or velocity decays after the initial program arc completes and the tree enters maintenance drift (ledger drift compounding, CI bit-rot, pandas 3.0 widening the parity gap). The base case leans toward the second shape on bus-factor-1-by-policy priors.

**60-month base case [Inference, Low]:** the software is either a niche artifact with a small `import frankenpandas as pd` following or dormant — the abandonment end state in the §4.7 steelman. The methodology-export (certified-lane benchmarking, negative-evidence ledgers, differential-against-oracle conformance, fail-closed runtime policy) is the durable survivor either way.

**Bifurcation — what the upside requires:** all four, in rough order: (a) the rider narrowed or removed (trigger 3) — without this nothing else matters for adoption; (b) a contributions on-ramp (trigger 4) — the policy refusal is the stranger ceiling; (c) a tagged release at a green pin with the parity numerals attested (triggers 1, 2); (d) an independent benchmark or review (trigger 6). The rider is the *necessary* condition, not merely first in a list [Inference, Low].

**Bifurcation — what decay looks like:** commit velocity falls off a cliff (trigger 9), the daily CI batch goes red or silent, the ledgers drift past the beads audit cadence, pandas 3.0 moves the specification, and the tree freezes as a 649k-line monument. Decay is the default outcome for bus-factor-1-by-policy infrastructure without institutional capture — the base case, not the tail.

**Revisit triggers (observable, falsifiable):**

1. **A tagged release or GitHub Release pointing at the assessed commit appears** — `git for-each-ref refs/tags` / `gh release list` non-empty *at the assessed commit*. (v0.3.0 exists — released 2026-09-12 — but targets the earlier tag commit `0ad22f91`, so this stays open.)
2. **The 119/119 and 1484/1484 numerals get a checked-in instrument** — a current `pandas_api_listing.json` (or equivalent) in-tree with the drift script green at the pin. Flips the headline parity claim from README-only to evidenced.
3. **Rider narrowed or removed** — LICENSE diff dropping the named-party restriction (and the destroy-copies clause). Flips usability for restricted parties and contributor-pool risk. v2/v3 note: bead `frankenpandas-dio8` (closed 2026-04-23) reviewed dropping the rider and the rider still stands at the pin — five months of post-review retention. Deliberate retention vs retention-by-default is adjudicated in §4.8 at [Inference, Medium]; either way, do not expect this trigger — treat any movement here as a genuine policy reversal [Code-verified, High on the bead; Inference, Medium on the expectation].
4. **A contributions on-ramp appears** — the "About Contributions" refusal amended, or a second human committer with merge rights observable on the contributor graph. Flips succession risk. (The stranger ceiling — it requires a policy reversal, not just a hire.)
5. **The shared thread pool lands and the 10 certified losses are re-measured** — the spawn-cost causal story confirmed or rejected per lane (next-step 5). Flips the concurrency pillar from story to evidence.
6. **Independent benchmark or review published** — any third party reproducing or contesting the certified-lane numbers. Flips "zero independent coverage."
7. **Conformance gate green at a pin** — the live-oracle parity job passing with the current DISCREPANCIES ledger. Flips the red-gate judgment.
8. **A Polars/pandas move collapses the differentiation** — Polars shipping a pandas-compat API layer, or pandas 3.0 absorbing the "auditable alignment" story — the kill-test: differentiation collapses to "Rust, but slower to adopt" [Inference, Medium].
9. **Abandonment tripwire** — no commits for 60 days or the CI workflows disabled/archived: re-ring to Monitor and treat the tree as a retired artifact.

## 4.12 Verdict & NODUS ring

**NODUS ring: Explore.** The Rulebook's Explore rule is explicit: *the default for substantive-but-unproven*. FrankenPandas is the textbook case: substantive (648,671 lines, a genuine zero-`unsafe` census, 1,387 differential packets, 8,954 tests, a 143-lane certified benchmark gate, a 44,086-line negative-evidence ledger — all [Counted/Code-verified, High]) and unproven (no release artifact at the pin, red parity gate, README-only parity numerals, zero independent coverage, bus factor 1 by policy, no production users). "When in doubt, ring down" does not apply — there is no doubt about the Explore floor; the doubt is about *evaluability by us*, which is not what the maturity scale measures. The tension the verdict must name: the ring says Explore while every advancement path is blocked — which reads as "Explore, but actually Monitor-in-practice." v2/v3 resolve it this way: the blocks are not transient. The rider was adopted at inception and survived explicit review without being dropped (bead `frankenpandas-dio8`, closed 2026-04-23; §4.8 adjudicates deliberate-retention vs retention-by-default at [Inference, Medium]); the contributions refusal is a stated policy, not an oversight. Both ceilings are therefore *durable* — the rider's "by choice" is inferred from sustained post-review tolerance rather than a recorded decision — and the ring logic prices them as durable rather than assuming their removal (see §4.9). The ring reads: **Explore — substantive, unproven, and durably ceilinged; the methodology is worth tracking, the software is not advanceable on any evidence-based trigger currently in view** [Inference, Medium]. No Rulebook amendment; v1.0 stands.

**The one-paragraph case:** FrankenPandas is a 649k-line bet that an entire API can be re-derived in safe Rust with machine-checked parity and published proof of every claim — and on process grounds it is winning that bet further than any peer: zero counted `unsafe`, 1,387 conformance packets against a pinned live oracle, ~9k tests, a benchmark gate that lists its own losses with thread counts, and a 44k-line ledger of its own failures including retracted wins. But the product case is hollow at the pin: the 4x number is a certified-subset geomean its own author discounts below 1.5x, the parity gate is red, the headline Python-parity numerals are README-only, one human is the entire bus factor *by stated policy*, the older bulk of the tree is agent-co-authored (which murkies the "clean-room" label), and the license forbids the likeliest evaluators from even analyzing the code — then orders destroyed copies on breach. Explore the methodology; the software stays behind the double ceiling until triggers 3 and 4 fire. [Inference, Medium — a recommendation plus a forecast, both the analyst's judgment.]

## 4.13 Limitations and open questions (analyst-facing)

**What was not done:** the workspace was never compiled; nothing was executed; no test suite was run; no benchmark was reproduced; the full `artifacts/bench` corpus was listed (960 rows) but row contents were only sampled; `Cargo.lock` was not in the sparse checkout (dependency posture partially inventoried); the fuzz corpus and `artifacts/phase2c` were not pulled; crates.io publication unverified (HTTP 403 from this network); CI job logs need admin rights (the conformance job's failing step is unestablished beyond its 1m5s fast-fail and the annotations); the full ~8,600-commit history was not pulled (authorship analysis covers ~700 commits via the API); review depth (who reviewed the 4,049 closed beads) unassessed; the Bedrock invariants 4–5 (zero panics, determinism) not verified by execution; demand-signal search performed 2026-09-22 (no independent coverage found beyond the repository).

**Open questions that would most change the verdict, in order of load-bearing weight:**

1. **What fraction of the tree's lines are agent-written?** Trailer counts bound commits touched, not authorship share — and the trailer convention lapsed for the newest ~300 commits. A `git blame` sample over full history would settle whether "clean-room" is a defensible label or a murky one.
2. **Does the rider bar this very assessment?** The license's "analyzing" prohibition arguably covers the FrankenSuite's evaluation activity — and the older bulk of the tree was co-authored by Anthropic models. A legal reading either way changes whether the program may continue touching the repo at all.
3. **Why is the conformance job red?** The live-oracle parity gate failed in 1m5s on the latest run — fast-fail suggests a harness/gate refusal rather than a full oracle pass. The failing step is unestablished (logs need admin rights).
4. **Why is fuzz-nightly red five runs straight, and what did the wheels failures at the v0.3.0 tag mean?** Five consecutive fuzz-nightly failures could be infra or found crashes; five wheels failures at the release commit sit oddly beside the green wheel job in CI.
5. **Are the 119/119 and 1484/1484 numerals reproducible?** Running the coverage-drift script against a pinned pandas 2.2.3 would attest or retire the README's most-repeated claim.
6. **Is there any production user or demand signal?** One real deployment would move the impact-potential score and the 24-month trajectory more than any code change. (Search 2026-09-22: none found.)
7. **What is the review depth?** Whether any of the 4,049 closed beads saw a second pair of eyes — determines whether the bead tracker is institutional knowledge or a solo diary [Not verified].
8. **What does the dependency tree actually contain?** `Cargo.lock` was not pulled; the unsafe-surface and license posture of the full transitive closure (PyO3, arrow-rs, parquet, rusqlite…) is unassessed.

---

## Deepening questions — Rulebook §5's binding eight, one paragraph each

*§5→location mapping: Provenance ← packet `fixture_provenance` + RaptorQ envelopes + EvidenceLedger (built on claims 7–8); Embeddable unit ← new, sourced from the crate topology (§4.4); Unexercised option value ← roadmap Medium/Low items + the asupersync example; Benchmark honesty ← §4.5 plus the NEGATIVE_EVIDENCE methodology rows; Governance path ← §4.8 plus §4.7 weakness 2 (bus factor by policy); License as strategy ← §4.8 plus §4.7 weakness 1 (rider); Agent-era fit ← new; Kill test ← §4.7's bear-case steelman plus trajectory trigger 8.*

**1. Provenance.** The repo records provenance at three layers, and v2/v3 verified the first by census, not sampling: `fixture_provenance` is present in **1,387 of 1,387** packet JSONs [Counted, High] — a sampled entry carries `{"pandas_version": "2.2.3", "oracle_script_sha256": "51158fcc…", "generated_at": "2026-09-01T06:08:59Z", "generation_command": "/data/tmp/frankenpandas-whiteotter-tn6qb2-build/debug/fp-conformance-cli …"}` — note the generation command's worker build path, consistent with the multi-host build story behind the 2.6x build-variance finding. RaptorQ envelopes wrap durable artifacts (fixture bundles, benchmark baselines, migration manifests) with a manifest, an integrity scrub report, and a decode proof per recovery event [Code-verified, High on the format]; and the EvidenceLedger logs every policy-driven decision with timestamps, mode, priors, posteriors, and Bayes factors [Code-verified, High on existence]. Portable attestation would require hash-chained, signed, write-time-bound entries — none of the three layers signs anything, and a hostile operator with disk access could rewrite a ledger undetectably [Inference, Medium]. What nothing records is *who* produced an artifact — the ledgers log events and decisions, not authors — so provenance-of-code (the §4.8 agent-authorship question: ~398 agent-co-authored commits in the sampled older window, unmeasurable in the newest ~300) is entirely unaddressed by the machinery [Inference, High].

**2. The embeddable unit.** The smallest useful piece adoptable without the whole repo is `fp-types` (Scalar/DType/NullKind/Timestamp as proper value types — the three-way null semantics in a box) or, one level up, `fp-columnar` (typed `Arc<[f64]>` backings + bitpacked `ValidityMask` + AG-10 vectorized kernels: a SIMD-without-unsafe columnar substrate) [Inference, Medium]. Adoption cost: every in-repo usage is a relative `path` dependency — v2 grepped `benches/` and `scripts/` and found **zero** external consumers of any `fp-*` crate, so the unit is vendored-only in practice, not just in principle — and it inherits the workspace's nightly toolchain (`portable_simd`), the 2024 edition, and the rider [Code-verified, High on the grep; Inference, Medium on the rest]. The hand-rolled pieces that *look* embeddable mostly aren't: the expression engine and the plotting renderer live inside `fp-frame`, not separate crates [Code-verified, High] — which concentrates the maintenance cliff on one human, exactly the §4.7.7 concern.

**3. Unexercised option value.** Three capabilities the architecture holds but has not used. First, lazy evaluation / query planning sits at roadmap "Medium" — the AACE alignment-planning phase is already an explicit plan-then-materialize split, so the plan object exists; only the cross-operation optimizer doesn't [Maintainer claim, High on the roadmap status]. Second — corrected in v2, which found v1's "single example" framing wrong: the asupersync interop is a **1,165-line, 7-file artifact-transfer subsystem** (codec with `ArtifactCodec`/`EncodedArtifact`/`PassthroughCodec`, capability-gated config, `Fnv1aVerifier` integrity proofs, recovery policies, an in-memory `TransportLayer` with `TransferReport`s), feature-gated off by default and consumed by `fp-conformance` behind the gate, plus `outcome_to_action` mapping the external `asupersync::Outcome` → `DecisionAction` (Ok→Allow, Err→Repair, Cancelled/Panicked→Reject) [Code-verified, High]. The unexercised option is therefore larger than v1 thought: a ready-made durable-artifact transfer layer the RaptorQ-envelope pipeline could actually ride on, currently held at arm's length behind a feature flag nobody turns on. Third, the shared thread pool: the 147-site `thread::scope` fan-out is the measured bottleneck (350–420µs/call), the fix is tracked, and landing it converts the 100k-row loss story from causal account to closed loop [Maintainer claim, High]. What unlocks each: an optimizer behind the alignment plan, a first real consumer of the asupersync transfer surface (or a decision to delete it — 1,165 lines of gated code with one consumer is also a liability), and the pool PR itself [Inference, Medium].

**4. Benchmark honesty.** The numbers most likely to survive an independent rerun are the large structural deltas (the O(n²)→O(n) sweep, the AXPY reorder, arena-backed groupby) — large enough to clear the maintainer's own build-variance bar [Inference, Medium]. The 3.967x certified-lane geomean would likely not reproduce as a number: v2 aggregated the funnel — 359 lanes decompose as FASTER 201 / SLOWER 20 / NULL_UNDECIDABLE 74 / DROPPED_HIGH_CV 62 / PARITY 2, with the certified 143 = 133 wins + 10 losses and 216 lanes uncertified — so the geomean is a *decidable-subset* statistic whose selection pressure runs toward lanes where the apparatus could reach a verdict; newest certified rows date to 2026-09-01 against a moved tree (STALE ELF flags); cross-worker build variance (measured 2.6x, same source) is uncontrolled in the corpus [Code-verified, High on the inputs; Inference, Medium on the conclusion]. The honesty apparatus itself is the load-bearing asset: per-arm A/A nulls, the bootstrap median-CI gate, ELF pinning with stale flags, certified losses listed with thread counts (including losses *with* a thread advantage), the fenced "FP-vs-FP, not vs pandas" optimization table, and the maintainer's standing rule that sub-1.5x ratios are UNRESOLVED [Code-verified, High]. No independent reproduction is known [External, High within recall caveats]; reproduction needs the pinned nightly toolchain, a pandas 2.2.3 oracle, and a single pinned host. The thesis rests on the parity evidence, not the speed numbers — the project's own invariant ordering says so [Inference, Medium].

**5. The governance path.** The credible route from one maintainer to an institution runs through four gates in rough order: rider removal, a contributions on-ramp (the policy refusal is the stranger gate — it requires a *reversal*, not just a hire), a tagged release at a green pin with the parity numerals attested, and an independent benchmark (trajectory triggers 3, 4, 1/2, 6). The starting position: no succession plan, no second committer, no foundation, no release artifact at the pin, and an explicit refusal to accept contributions [Inference, Medium on the absences — the no-release half is Git-observed, High]. The closest thing to institutional knowledge transfer is the in-repo bead tracker (4,051 records, 4,049 closed [Counted, High]) plus the evidence ledgers — process artifacts that outlive a maintainer, and the only onboarding story a successor would inherit; review depth is unassessed [Not verified] and parked as §4.13 open question 7. What breaks first if velocity decays: the daily CI batch — the fast-gate's own header documents a 4-day unnoticed uncompilable stretch — then ledger drift; the 60-day abandonment tripwire (trigger 9) is the observable sensor [Inference, Medium]. The realistic institutional endpoints are a foundation home or a corporate adopter — both currently gated on rider removal *and* the policy reversal; v2 reframes these not as advancement blockers awaiting removal but as durable ceilings by the maintainer's choice (the rider was explicitly reviewed and retained in bead `frankenpandas-dio8`), which is why the ring logic in §4.9/§4.12 prices them rather than assuming them away [Inference, Low].

**6. The license as strategy.** The rider excludes exactly: OpenAI, L.L.C.; Anthropic, PBC; their affiliates; and anyone "acting directly or indirectly on behalf of, for the benefit of, or under the direction of" them — barring "use" defined expansively (copying … benchmarking, testing, **analyzing, indexing** … incorporating into datasets, training corpora, evaluation harnesses, or ML pipelines), with automatic termination plus **destroy-all-copies** on breach and injunctive relief reserved [Code-verified (license text), High]. Does the exclusion serve or sabotage the stated mission? The committed answer, revised in v2 with the maintainer's own policy record: **it serves its defensive intent and sabotages the stated mission — and the sabotage is the binding constraint.** The defensive intent is no longer "unlabeled," as v1 claimed: bead `frankenpandas-dio8` (closed 2026-04-23) framed the rider as "intentional policy against certain AI providers' training-data use," offered dual-licensing as an alternative, and closed with the rider intact — though its close reason defers rather than records a choice (deliberate retention vs retention-by-default adjudicated in §4.8 at [Inference, Medium]) — and the §4.7.7 finding (the older bulk of the tree heavily Claude-co-authored) makes barring Anthropic from training on it at least a coherent motive [Code-verified, High on the bead; Inference, Medium on the motive]. But the mission as stated in the README is *auditability and evidence* — parity packets, certified lanes, published retractions — and the rider taxes exactly that currency: the entities with the resources to validate a 649k-line safety claim may not benchmark, test, analyze, or index the code, while the "acting for" prong creates legal uncertainty for contributors nowhere near the AI supply chain. A defensive win that costs the mission its evaluability is a net loss, because evaluability is the mission [Inference, Medium]. The remaining strategy point stands: the rider is a licensee-facing contract term — it binds users, not the author — so it neither creates nor answers any copyright question about the agent-co-authored tree; its strategic function is access control, not clean-room hygiene [Inference, Medium].

**7. Agent-era fit.** The concrete agent workload that would pick this over the incumbent is the pandas-code-execution workload: an agent team running untrusted or legacy pandas snippets could swap `import pandas` for `import frankenpandas as pd` and get memory safety, no GIL, deterministic seeds, and an EvidenceLedger audit trail of every alignment and coercion decision the run made — the fail-closed modes (Strict rejects, Hardened logs-and-recovers under a Bayesian loss rule) are exactly the shape an agent-safety review wants around data pipelines [Inference, Medium]. What would have to become true first: the rider narrowed or removed (AI labs and their agents are currently barred from even *analyzing* the code — the target user is excluded by name), a contributions on-ramp or at least a support story (there is literally no one to file a blocking issue *with* — issues are welcome, fixes are at the maintainer's discretion), the parity numerals attested (the "drop-in" claim is README-only), and the conformance gate green (it is red). The fit is real but presently blocked by the license aimed at its own likeliest users — and the recursion is noted in §4.8 (this packet is itself an agent-assisted assessment of agent-co-authored code under an anti-agent-analysis rider) [Inference, Medium].

**8. The kill test.** The single experiment that would falsify the core thesis — that a clean-room Rust pandas with machine-checked parity and a published evidence discipline is a differentiable, adoptable artifact — is the one the project built for itself: re-run the certified census on a single pinned host at a new pin. If the 3.967x certified-lane geomean collapses (below ~2x, say) once STALE ELFs are re-measured and build variance is controlled, the performance pillar is falsified by the project's own instrument, not by a critic [Inference, Medium]. Two secondary falsifiers sit behind it: a competitor move (Polars shipping a pandas-compat API layer, or pandas 3.0 absorbing the auditable-alignment story — trajectory trigger 8, collapsing the differentiation to "Rust, but slower to adopt") and the abandonment tripwire (trigger 9: 60 days without commits or disabled CI re-rings to Monitor). The production checklist is the inverse of this test — what success requires, scored at the pin as 0 of 8 fully met: tagged release at the assessed commit, attested parity numerals, green conformance gate, the shared-pool fix with re-measured losses, closed parity-tail DISCs (011/006/009), a signable license, a contributions on-ramp, bus factor > 1 [Inference, Medium].
