# FrankenLibC — Research Assessment Packet

**Repository:** [Dicklesworthstone/frankenlibc](https://github.com/Dicklesworthstone/frankenlibc) · **Language:** Rust (nightly) · **License:** MIT + OpenAI/Anthropic rider (non-OSI; see §4.8)

**Pinned commit:** `36e5be9448c8c26eed16f421672775aae048f936` — *"feat(resolv): complete getnameinfo native reverse DNS and service lookup"*, 2026-09-22 15:54 UTC

**Stars:** 50 · **Forks:** 3 · **Watchers:** 1 · **Last push:** 2026-09-22 · **Releases:** none · **Tags:** none

**Assessment date:** 2026-09-22 · **Assessor:** Josh's research program (coordinator packet)

**Method:** fresh blobless clone at the pin, live GitHub Actions observation, full code/artifact inspection, competitor and coverage search. **Not done:** no build executed, no benchmarks reproduced, no fuzzers run, no runtime-math kernel audited for mathematical validity.

---

**Hook:** A 2.3-million-line Rust reimplementation of glibc that runs real binaries under `LD_PRELOAD` — except a million of those lines are generated iconv tables, a third of its symbols still phone home to the host glibc, and its license forbids Anthropic's models from analyzing code that Anthropic's models helped write.

---

**Tier legend (Rulebook §1):** **[Verified]** direct inspection of the pinned clone or a live page read by the analyst — flavors **[Counted]** (I ran the count), **[Git-observed]** (git metadata), **[Code-verified]** (source read). **[CI-observed]** is Tier 2 (seen executing on live CI pages — attests the suite *runs*, not that it is green). **[Maintainer claim]** asserted in README/docs, not independently executed. **[External]** independent sources. **[Inference]** analyst judgment, always labeled. Confidence: **High** / **Medium** / **Low**.

## TL;DR

FrankenLibC is a Linux `LD_PRELOAD` interposition layer written in Rust that reimplements the C library ABI surface: 4,119 exported symbols individually classified (69.3% native — own Rust implementation or direct syscall; 30.7% validated call-through to host glibc), run through a policy membrane with strict/hardened runtime modes, an optional 71-kernel runtime-math control plane, 66 fuzz targets, and an unusually honest claim/evidence taxonomy that distinguishes competitive wins from self-speedups. **Verdict: TRL 4 · NODUS Explore.** Its most important strength is the claim/evidence machinery — the 41,205-line negative-evidence ledger, same-invocation incumbent benchmarks, and published self-retractions are the strongest measurement discipline observed in the program so far [Inference, High]. Its most important ceiling is that main CI is red at the pin, the shipping artifact is still L1 interposition over host glibc, the bus factor is one, and the MIT+Anthropic/OpenAI rider makes independent evaluation by the best-equipped labs a license violation [High].

---

## Quick Links

All links verified resolving on 2026-09-22 (at the pinned commit unless noted).

1. [Repository root](https://github.com/Dicklesworthstone/frankenlibc)
2. [README at the pin](https://github.com/Dicklesworthstone/frankenlibc/blob/36e5be9448c8c26eed16f421672775aae048f936/README.md) — 4,020 lines / 247 KB
3. [LICENSE at the pin](https://github.com/Dicklesworthstone/frankenlibc/blob/36e5be9448c8c26eed16f421672775aae048f936/LICENSE) — read verbatim
4. [Symbol support matrix at the pin](https://github.com/Dicklesworthstone/frankenlibc/blob/36e5be9448c8c26eed16f421672775aae048f936/support_matrix.json) — 4,119 rows / 983 KB
5. [Negative-evidence ledger at the pin](https://github.com/Dicklesworthstone/frankenlibc/blob/36e5be9448c8c26eed16f421672775aae048f936/docs/NEGATIVE_EVIDENCE.md) — 41,205 lines / 3.39 MB
6. [Current performance evidence at the pin](https://github.com/Dicklesworthstone/frankenlibc/blob/36e5be9448c8c26eed16f421672775aae048f936/docs/PERF_FRONTIER_FINAL.md) — dated 2026-07-27, campaign-win taxonomy
7. [Actions — `ci.yml`](https://github.com/Dicklesworthstone/frankenlibc/actions/workflows/ci.yml) — pin's run observed red
8. [Agent operating manual (AGENTS.md) at the pin](https://github.com/Dicklesworthstone/frankenlibc/blob/36e5be9448c8c26eed16f421672775aae048f936/AGENTS.md) — 2,109 lines; contains the silent-deletion audit
9. [Deployment guide at the pin](https://github.com/Dicklesworthstone/frankenlibc/blob/36e5be9448c8c26eed16f421672775aae048f936/DEPLOYMENT.md) — honest scope boundaries quoted in §4.2
10. [Smoke-battery summary at the pin](https://github.com/Dicklesworthstone/frankenlibc/blob/36e5be9448c8c26eed16f421672775aae048f936/tests/conformance/ld_preload_smoke_summary.v1.json) — checked artifact, 2026-06-03

## Did You Know?

FrankenLibC's agent operating manual *mandates* a 44-family "Mandatory Modern Math Stack" — every major subsystem milestone "must use at least 3 distinct math families", each milestone "must include at least one obligation from conformal statistics, algebraic topology, abstract algebra, and Grothendieck-Serre methods", and "no single family should dominate more than 40% of the milestone obligations" (AGENTS.md, read verbatim at the pin). And the README's FAQ asks its own hardest question outright: **"Is the runtime math real code or just naming theater?"** The maintainer's answer points to the source files and a kill-switch. The kill-switch is real: `FRANKENLIBC_RUNTIME_MATH=off` exists so operators can disable the whole 71-kernel control plane without rebuilding. [Code-verified — High]

## Franken-Worthy Next Steps

1. **Import the campaign-win/self-speedup taxonomy into the FrankenSuite program protocol.** FrankenLibC's rule — a win needs the incumbent live *in the same invocation*; everything else is maintenance — is the strongest anti-benchmark-theater device observed in the program so far [Inference, High]. Falsifiable: check the program protocol document for a machine-readable `result_class` field on every performance claim; currently absent in all sibling packets reviewed.
2. **Run the per-commit deletion-excess audit on one sibling agent-assisted repo.** FrankenLibC's method (author-date filter, rank by deletions−insertions, adjudicate against the ledger) is a reproducible screen for silent work destruction. Falsifiable: a one-page writeup naming the repo, the query, and the adjudicated verdicts, or a finding of zero deletions.
3. **Extract the claim-field contract as a reusable honesty schema.** `symbol_status` ≠ `semantic_parity` ≠ `replacement_level` ≠ `freshness_state` — separating these four killed at least three classes of misleading claim in one repo. Falsifiable: a schema document plus one other packet in the suite rewritten against it without information loss.
4. **Second-party verification of the bead-closure audit.** The project admits four beads closed on gates that could not link ("an unlinkable target is silent, not green"). Sample the 7,422 closed beads for the "compilation observed" rule and publish the hit rate. Falsifiable: a published sample with the count, which the maintainer's own audit predicts is non-zero.
5. **A/B the runtime math with its own kill-switch.** Run the conformance/smoke battery and the frontier benchmarks with `FRANKENLIBC_RUNTIME_MATH=off` versus on, and publish what the 71 kernels actually buy. Falsifiable: a signed artifact with the deltas; a null result answers the README's "naming theater" question against the maintainer.

---
## 4. The Full Assessment

### 4.1 Why it exists — the market problem

A C standard library is the most load-bearing software nobody can replace. glibc is ~30 years of accumulated behavior: undefined-behavior corners, locale quirks, allocator habits, and extension APIs that the entire Linux userspace implicitly depends on. The demand side of this problem is real and well documented: memory-unsafety bugs in C libraries remain a top source of CVEs, and the CISA-led memory-safety movement is pushing the industry to stop writing new memory-unsafe code. But rewriting a libc is a different sport from rewriting a service. The ABI surface is vast (thousands of symbols), behavior must be bug-for-bug compatible with programs that were never written correctly, and the failure mode of getting it wrong is *everything* crashing — so nobody adopts a replacement until it is proven, and it cannot be proven without being adopted.

FrankenLibC's answer to that trap is **interposition instead of replacement**. Rather than asking the world to swap `libc.so.6`, it ships `libfrankenlibc_abi.so` and asks to be `LD_PRELOAD`ed *on top of* host glibc: intercept every call, validate the inputs through a policy "membrane", serve the implemented ones from Rust, and call through to the host for the rest. The replacement-levels model in the repo is explicit about staging this: L0 (interpose, call-throughs allowed), L1 (hardened interpose — the current release claim), L2 (standalone `.so`), L3 (full replacement). This is a legitimate staging strategy for an unreplaceable component — but it also caps every deployment claim: the shipping model "still depends on host glibc" (DEPLOYMENT.md, maintainer's own words), and LD_PRELOAD cannot touch static binaries, setuid/setgid binaries, or code that runs before the preload resolves. The alternatives clarify the bet: improving musl concedes the glibc-ABI compatibility that *is* the problem; sanitizer-style instrumentation needs the world recompiled; seccomp/eBPF can police syscalls but cannot validate libc-level semantics (locale behavior, allocator invariants, string-function contracts). Interposition is the only approach that validates the library boundary without rebuilding the world — which is also why its ceiling is structural: the validator must always run inside the process it distrusts.

**Demand-side check.** A web search for FrankenLibC found no independent coverage: no benchmarks, reviews, deployments, blog posts, or downstream dependents — only the repository, two named forks, and its renamed predecessor (`glibc_rust`, same account, now redirecting to the current name). Search was English-language web; non-indexed or non-English discussion cannot be ruled out. Demand is therefore maintainer-asserted, not market-observed. The forks' most recent pushes could not be verified from the API data in hand, and the contributor graph shows a single account. [Inference — no market validation exists either way]

**Why now.** Three conditions make this attempt legible: Rust's mature enough to write this at all (nightly 2026 features, 2024 edition); the industry's memory-safety pressure creates the only plausible adoption motive; and AI-assisted coding velocity makes a ~50-commit-per-day pace on a 4,119-symbol surface possible for a solo maintainer — which this repo demonstrates at scale.

**Adjacency.** FrankenLibC is one of 44 "franken" clean-room Rust reimplementation repos run by the same maintainer (the "FrankenSuite"). Shared tooling (`asupersync-conformance`) is used across the suite as dev/test orchestration, not runtime code — verified in §4.4, correcting earlier draft language that overstated the dependency.

### 4.2 What it is — repository TL;DR and one-line verdict

FrankenLibC is a Rust (nightly-2026-08-31, edition 2024) reimplementation of the Linux C library ABI, shipped as `frankenlibc-abi` (`cdylib`/`staticlib`/`rlib`) and deployed today by `LD_PRELOAD`ing it over host glibc. Its distinguishing machinery is not the libc surface itself but the *claim infrastructure around it*: every one of 4,119 exported symbols carries a machine-checked support status; behavior claims are separated into symbol-presence, semantic-parity, replacement-level, and evidence-freshness fields; benchmark rules require the incumbent glibc to run live in the same process invocation for a "campaign win" (everything else is a "maintenance self-speedup"); and a 41,205-line negative-evidence ledger records rejections, losses, and retractions — including retractions of the project's own stale wins.

**One-line verdict:** The most rigorously self-audited C-library rewrite this program has assessed [Inference, High], and still an L1 interposition prototype whose strongest proven asset is its evidence machinery — which is why the ring is Explore, not Monitor.

### 4.3 Claim inventory

Each claim is scored with a Rulebook status — **Demonstrated** (independently checkable from the pinned tree or live observation), **Partially demonstrated** (real evidence, material gap), **Aspirational** (roadmap), **Disproven** (contradicted by the evidence), **Stale** (code moved past the docs) — plus a confidence grade.

| # | Claim | Status | Basis |
|---|-------|--------|-------|
| 1 | 4,119 exported symbols, each classified: 2,441 `Implemented`, 414 `RawSyscall`, 1,264 `WrapsHostLibc` | Demonstrated | Counted from `support_matrix.json` at the pin — [High] |
| 2 | 69.3% of symbols served natively (own Rust or direct syscall), 30.7% wrapped host glibc | Demonstrated | Same count. Note "native" is the maintainer's taxonomy term, not an independence claim — the membrane still validates the wrapped calls — [High] |
| 3 | Shipping artifact `libfrankenlibc_abi.so` works via `LD_PRELOAD` today; current claim L1 | Demonstrated | Deployment model confirmed by build artifact name, smoke battery, and DEPLOYMENT.md; L1 stated by maintainer — [High] |
| 4 | 66 cargo-fuzz targets exist | Demonstrated | Counted 66 `fuzz_*.rs` targets — [High] |
| 5 | Codebase is ~2.3M Rust lines across 2,697 files (source-only ~1.45M; ~452k excluding generated iconv tables) | Demonstrated | Fresh counts at pin; v1's "~304k LOC" figure is disproven as a methodology artifact — [High] |
| 6 | Main CI passes at the pinned commit | Disproven | Pin's main CI run #8942: Core Gates failed on its first step (`scripts/ci.sh`); all substantive jobs skipped. A separate candidate-validation workflow for the pin's change was green, but that is not main CI — [High] |
| 7 | Smoke battery 64 cases: 60 pass, 0 fail, 4 optional skips | Partially demonstrated | Checked-in artifact is real and explicit (2026-06-03), but 3.5 months stale; README discloses the date — [High] |
| 8 | Strict and hardened runtime modes are live | Partially demonstrated | Mode machinery, env-var contract, and kill-switches are in the tree; behavioral evidence is the June smoke artifact only — [Medium] |
| 9 | Campaign benchmark wins vs live glibc: `strftime` 0.54×, `wcsftime` 0.11× with bootstrap CIs | Partially demonstrated | Published with methodology (dlmopen-protected incumbent, self-comparison null controls) and dated 2026-07-27; independent reproduction is the gap — [Medium] |
| 10 | Published losses: `malloc_free` 5.64× slower than glibc (Sept 19); frontier sweep reports mixed results | Demonstrated | Ledger entries show the loss rows verbatim, with improvement from 6.87× — [High] |
| 11 | 71 runtime-math kernels operational | Partially demonstrated | Exactly 71 `.rs` files and 72 `pub mod` entries — existence verified; production value is asserted, and a runtime kill-switch ships for it — [High on existence / Low on value] |
| 12 | Unsafe code is tightly governed (`deny` in core/membrane, `forbid` in harness) | Partially demonstrated | Gates are real at crate roots, but the workspace lint allows unsafe by default and the ABI crate — the 4,119-entrypoint surface — explicitly allows it; 382 files carry SAFETY comments. Governance is process, not proof — [High] |
| 13 | Runtime depends on `asupersync-conformance` | Disproven | Verified: dev-dependency in membrane, optional/feature-gated in fixture/harness/conformance. `frankenlibc-abi` runtime deps are membrane, core, libc, parking_lot, sha2, md-5, tracing(optional) — [High] |
| 14 | Standalone `libfrankenlibc_replace.so` (L2/L3) exists | Aspirational | "Not shipped yet" — DEPLOYMENT.md — [High] |
| 15 | "Drop-in replacement for glibc" today | Disproven | Maintainer's own FAQ: "not a drop-in replacement for glibc today"; DEPLOYMENT.md warns not to repoint `libc.so.6` at it — [High] |
| 16 | Bus factor one | Demonstrated | Sole contributor account across 11,272 commits — [High] |
| 17 | ~49% of recent commits carry AI co-authorship (434 Claude-family, 3 Grok of 895 sampled, 2026-08-18→2026-09-22) | Demonstrated | Counted from commit trailers in a blobless-history sample; the pin itself has no trailer — [High] |
| 18 | Silent-deletion self-audit published; three stale banked wins retracted 2026-09-01 | Demonstrated | Documented in AGENTS.md with commit hashes, dates, and the retraction commits; ledger carries CORRECTION blocks — [High] |
| 19 | Runtime-math kill-switch `FRANKENLIBC_RUNTIME_MATH=off` ships | Demonstrated | In DEPLOYMENT.md env-var contract — [High] |
| 20 | Nightly-only toolchain, Rust 2024 edition | Demonstrated | `rust-toolchain.toml` pins `nightly-2026-08-31` — [High] |

---
### 4.4 Codebase tour

**Reconstructed architecture (from the pinned tree, not the diagrams).**

```
┌─────────────────────────────────────────────────────────┐
│ frankenlibc-abi (cdylib/staticlib/rlib)                 │  shipping artifact: libfrankenlibc_abi.so
│  C ABI entrypoints (libc.map: 4,749 lines)               │
│  explicit unsafe allow — the 4,119-symbol surface        │
└──────────────┬──────────────────────────────────────────┘
               │ calls through
┌──────────────▼──────────────────────────────────────────┐
│ frankenlibc-membrane                                    │  #![deny(unsafe_code)]
│  runtime_policy — per-call policy engine                │
│  runtime_math/ — 71 kernel files, "44 math families"    │
│  arena, fingerprint (scoped unsafe allow + SAFETY)      │
│  strict / hardened modes (FRANKENLIBC_MODE)             │
└──────────────┬──────────────────────────────────────────┘
               │ validates, then
┌──────────────▼──────────────────────────────────────────┐
│ frankenlibc-core                                        │  #![deny(unsafe_code)]
│  owned implementations (string, stdio, time, locale,    │
│  math, iconv tables, …)                                 │
└──────────────┬──────────────────────────────────────────┘
               │ or falls through to
┌──────────────▼──────────────────────────────────────────┐
│ host glibc — WrapsHostLibc path (1,264 symbols)          │
│ kernel direct — RawSyscall path (414 symbols)           │
└─────────────────────────────────────────────────────────┘
```

Around this core: `frankenlibc-fixture-exec` (fixture execution), `frankenlibc-harness` (`#![forbid(unsafe_code)]` — test/bench orchestration), `frankenlibc-bench`, legacy `frankenlibc` and `frankenlibc_conformance` crates, and `frankenlibc-fuzz` — a standalone cargo-fuzz crate (66 targets) that is **not** a workspace member, depending on membrane/core/abi by path. The root workspace lists 8 members; there is no checked-in `Cargo.lock` in the tree observed.

**Line counts, honestly reported.** Raw totals under `crates/`: 2,697 Rust files, 2,303,168 lines. Two methodologies are used in this packet and they must not be confused: the *per-crate table* below counts every `.rs` file outside `target/` (2.29M lines total — this includes tests, benches, and fuzz targets); the *"source-only ~1.45M"* figure counts only files under `*/src/` directories (393 files), excluding test/bench/fuzz-target trees. Both are stated so no single LOC number misleads. The distribution is extreme under either count: 24 files in `core/src/i18n/iconv/` carry ~1.0M lines of generated encoding tables (`cjk_tables.rs` alone is 260,589 lines), and `abi/tests/conformance_diff_fromfp.rs` is a 124,102-line generated differential test. Excluding iconv from the src-only count: 368 files, 451,742 lines. Any single "LOC" number for this repo is a methodology choice; the honest statement is that roughly two-thirds of the counted lines are generated tables and generated tests. A prior draft's "~304k" figure is superseded.

**Per-crate totals** (all `.rs` files excluding `target/`, counted at the pin): core 209 files / 1,116,280 lines (dominated by the iconv tables), abi 979 / 575,915 (dominated by the generated fromfp differential test and entrypoint glue), harness 1,069 / 372,461, membrane 126 / 88,477, bench 237 / 81,089, conformance 3 / 37,563, fuzz 66 / 20,587 (one file per fuzz target), legacy frankenlibc 6 / 632, fixture-exec 2 / 72. Note the last: `frankenlibc-fixture-exec` is a near-empty shim — the real fixture-execution machinery lives in the harness, a small piece of workspace untidiness worth knowing before trusting crate names. The membrane — the policy layer that justifies the project — is 88.5k lines: large for a policy engine, small relative to the 4,119-entrypoint surface it guards.

**Unsafe distribution.** The gates are real but scoped: `frankenlibc-core` and `frankenlibc-membrane` deny unsafe at the crate root while explicitly allowing it in specific modules (`arena`, `fingerprint`) with `SAFETY` comments; the harness forbids it outright; the ABI crate — the one that actually exposes all 4,119 entrypoints — explicitly allows it, which is unsurprising for FFI glue but means the *boundary layer itself* is outside the deny regime. The workspace lint defaults `unsafe_code = "allow"`, so the deny/forbid attributes are opt-in discipline per crate, not a workspace guarantee. Rough regex counts of `unsafe` tokens were deliberately not published: they count declarations and generated bindings, not unsafe sites, and would mislead. The correct characterization is governance-by-convention with documented exceptions — a process control, not a memory-safety proof.

**Dependency posture.** `asupersync-conformance = 0.5.0` is workspace-pinned but verified to be tooling-only: a `[dev-dependencies]` entry in membrane, and optional/feature-gated in fixture-exec, harness, and the legacy conformance crate. It is not in `frankenlibc-abi`'s runtime dependency closure (membrane, core, `libc`, `parking_lot`, `sha2`, `md-5`, optional `tracing`). The AGENTS.md "companion crates" table describes `/dp/asupersync` and `/dp/frankentui` as build/test tooling — "NOT runtime libc dependencies" — which matches the manifests. Earlier draft language implying a runtime dependency was wrong and is corrected here.

**Notable subsystems.** The runtime-math control plane (`membrane/src/runtime_math/`, 71 files, 12 design docs) implements per-call optimization policies under the "44 math families" doctrine mandated by AGENTS.md, with the branch-diversity rule (at least 3 distinct math families per major milestone; no single family to dominate more than 40% of a milestone's obligations). It ships with a documented kill-switch (`FRANKENLIBC_RUNTIME_MATH=off`, bead bd-06bxm.9) that skips kernel consultation while keeping basic membrane validation — a kill-switch for your own headline innovation is either admirable caution or a tell, and the packet does not resolve which. The Gentoo/Portage integration (`data/gentoo`, `docker/gentoo`, phase allowlists, package blocklists including `sys-libs/glibc` and `sys-apps/shadow`) is the only deployment lane with per-package operational controls. The bead tracker (`.beads/issues.jsonl`: 7,422 closed, 32 open) is the project's working memory — and, given the bus factor, most of its institutional knowledge. 588 `check_*.sh` gate scripts encode the claim/evidence contracts in executable form.

### 4.5 Maintainer's stated case — and a benchmark audit

**The pitch, fairly stated.** The maintainer's case is that a libc can be made safer and more honest without being replaced first: interpose over glibc, validate every call's inputs through a policy membrane, reimplement the hottest and most dangerous surfaces natively, and — the genuinely novel part — attach machine-checked claim metadata to everything so that no statement about the project outruns its evidence. The README's claim-field contract is explicit: *symbol status is not semantic parity, which is not replacement level, which is not evidence freshness.* That single discipline, if the industry adopted it, would end most benchmark theater.

**The benchmark rules are the best observed in the program so far [Inference, High].** The taxonomy (PERF_FRONTIER_FINAL.md, 2026-07-27): a **campaign win** requires FrankenLibC versus the actual host-glibc incumbent, run side-by-side *in the same invocation*, with the host arm protected from symbol interposition via `dlmopen(LM_ID_NEWLM)` and the ratio clearing a bootstrap median-CI/null-control gate. A **maintenance self-speedup** (FrankenLibC before vs after) "can justify a code change, but it is not a competitive claim or campaign output." The repo's own AGENTS.md lists twelve named reward-hacking patterns as forbidden (gate self-weakening, proof-class inflation, bench-path hardcoding, conformance metastasis among them) and states the three load-bearing rules: a self-speedup is maintenance not a win; never weaken a gate to land a change; reporting a loss is a success. This is the standard the project holds itself to — and then it publishes the losses anyway: the ledger's September 19 frontier entries show `getauxval` at 0.748× — a certified win (NEGATIVE_EVIDENCE.md L4403) — against `malloc_free` small-64 at **5.644× slower** than glibc, "still the largest certified family loss" though improved from the 6.87× headline of 2026-08-17 (L41204), and `thrd_current` at 2.278864× slower with a tight bootstrap CI (L32165). The two banked campaign wins (C-locale `strftime` 0.5406 [0.5371, 0.5496]; `wcsftime` alias 0.1113 [0.1069, 0.1164]) carry exact-fit coverage and 200,000-case differentials with zero divergences. That is honest measurement craft.

**Benchmark provenance table.** Every number below is maintainer-measured; none has been independently reproduced. Reproduction cost, stated honestly: an ~8,808-file checkout, a Linux host, the pinned `nightly-2026-08-31` toolchain, and compute for the frontier scripts — the smoke battery alone is 64 cases at 10-second timeouts plus build time. Not attempted here.

| Claim | Maintainer-measured | Independently reproduced | Standing |
|---|---|---|---|
| C-locale `strftime` 0.5406× vs live glibc [0.5371, 0.5496] | Yes — PERF_FRONTIER_FINAL.md, 2026-07-27, dlmopen-protected incumbent, A/A null control (a control run comparing the harness against itself), 200k-case differential, zero divergences | No | Banked campaign win |
| `wcsftime` alias 0.1113× vs live glibc [0.1069, 0.1164] | Yes — same methodology, 11 live-glibc conformance cases | No | Banked campaign win |
| `getauxval` 0.748× (win), `malloc_free` 5.644× slower, `thrd_current` 2.279× slower | Yes — NEGATIVE_EVIDENCE.md frontier entries, 2026-09-19 | No | Current frontier: mixed |
| Smoke battery 60/64 pass | Yes — checked-in artifact, 2026-06-03 | No (artifact read, not rerun) | Stale 3.5 months |

**The audit, and its limits.** Reproduction cost is the ceiling: no independent rerun was performed here, the frontier summary is dated 2026-07-27 (nearly two months before the pin), and the ledger's September 19 entries show a mixed, still-incomplete frontier — several rows report errors or missing incumbent arms. The methodology is demonstrated; the *current* performance picture is a snapshot in motion. And the project's own history supplies the sharpest caveat: three banked benchmark wins stood falsely in the ledger for ten weeks after their implementations were silently deleted, retracted only on 2026-09-01. The gates did not prove the mechanism existed. The project found this itself and published the correction — which is the strongest evidence both for and against trusting the machinery.

---
### 4.6 Competitors

| Competitor | Lane | Why it wins today | FrankenLibC's answer |
|---|---|---|---|
| **glibc** | The commodity incumbent | It is the ABI. Every Linux binary is built against its behavior, quirks included. Replacing it means re-proving the world. | Don't replace it — interpose over it. Correct as a staging strategy; concedes the endgame. |
| **musl** | The small, auditable production libc | Actually deployed (Alpine, containers, embedded). Clean codebase, static-linking story, real users. Owns "simple alternative libc" outright. | musl is clean C, not memory-safe Rust, and has no per-call validation story. Different lane — but musl owns the "trust this libc" narrative in production, and FrankenLibC has no production story at all. |
| **relibc** (pop-os/relibc) | The Rust libc peer | A Rust C/POSIX library, actively developed, targeting Redox and Linux. The closest direct peer in language and ambition. | relibc targets being a real libc for a real OS; FrankenLibC targets validating calls on Linux via interposition. The approaches barely overlap — which also means FrankenLibC's evidence machinery has no peer to be compared against. |
| **GrapheneOS hardened_malloc** | Hardened allocator, preloadable | Deployed at scale (Android/GrapheneOS), integrates as libc allocator or `LD_PRELOAD`. Documents the same interposition limits FrankenLibC faces (`AT_SECURE`/setuid). | Narrower (allocator only) but *proven in production*. FrankenLibC's whole-surface ambition is unproven; hardened_malloc is the existence proof that the preload lane can work and the warning label for its limits. |
| **Sanitizers / MTE / hardware tagging** | Adjacent memory-safety tooling | ASan/MSan need recompilation; MTE needs ARM hardware. Both are real, deployed, and understood. | FrankenLibC needs neither recompile nor new hardware — its one structural advantage. Untested at scale. |

**Why the incumbent wins.** glibc wins because it is not a product but a compatibility surface: three decades of accumulated behavior — including bugs that userspace depends on — are baked into every Linux binary ever shipped, and no replacement can be adopted until it reproduces that behavior bug-for-bug while no one can prove the reproduction without adoption. musl wins the "clean alternative" lane on the same logic in reverse: it is small enough to audit and already deployed (Alpine, containers), so it owns every conversation about replacing glibc that does not require glibc's quirks. FrankenLibC's interposition strategy is the only approach in the table that sidesteps this trap rather than charging it — which is why its lane is empty, and why the lane's emptiness proves nothing about demand. [Inference]

**The unoccupied lane.** Nobody currently owns *whole-libc ABI interposition with per-call validation and machine-checked claim metadata*. Sanitizers need recompiles, hardened_malloc covers only allocation, musl is a replacement not a validator, relibc is an OS libc. If the thesis is "validate the boundary you cannot replace," the lane is empty — but it is empty of evidence too: no independent deployment, no third-party benchmark, no user. An empty lane with no foot traffic is not a market; it is a hypothesis. The falsifier to watch: if a serious security team evaluates L1 interposition for a hardened fleet and rejects it on overhead or coverage grounds, the lane thesis weakens more than any competitor could weaken it.

### 4.7 Technical merit and adversarial review

**Three substantive strengths.**

**1. The claim/evidence machinery is the best observed in the program so far [Inference, High].** Every one of 4,119 symbols carries a machine-checked support status; symbol status, semantic parity, replacement level, and evidence freshness are separate fields so no claim outruns its evidence; 588 executable gate scripts enforce the contracts in CI; and the 41,205-line negative-evidence ledger records rejections, losses, and dated CORRECTION blocks. This is infrastructure for honesty, and it is the project's most portable asset.

**2. The benchmark methodology is genuinely adversarial-grade.** The campaign-win/self-speedup distinction, the same-invocation incumbent requirement with `dlmopen` protection, bootstrap median CIs, A/A null controls, and twelve named forbidden reward-hacking patterns constitute a measurement regime most industry benchmarks fail. The project then publishes its losses (malloc 5.644× slower) under the same regime.

**3. The self-audit culture is real, not decorative.** The silent-deletion audit names commits, dates, line counts, and the maintainer's own failures; the bead-closure rules ("a bead may not close on a gate whose compilation was not observed") were written from four observed violations; the README discloses the smoke artifact's date and calls it "a curated workload signal, not broad production workload readiness." A project that documents its own stale wins retracting is rarer than a project with no stale wins.

**The skeptic's take — eight weaknesses, tiered by severity.**

**[FATAL] The rider poisons the well it needs most.** The MIT+Anthropic/OpenAI rider (quoted verbatim in §4.8) forbids the named labs — and anyone acting for their benefit — from benchmarking, testing, analyzing, or indexing the code. 48.8% of recent commits carry AI co-authorship, 434 of them Claude-family. The project's velocity engine is substantially Anthropic models; its license forbids Anthropic from touching the output. Independent evaluation by the organizations best equipped to audit a 4,119-symbol Rust libc is a license violation. No amendment or workaround was found in the tree.

**[HIGH] The safety kernel is unsafe code wearing a deny gate.** `deny(unsafe_code)` at the crate root of core and membrane coexists with explicit `allow` in `arena` and `fingerprint` — the allocator and the validation-fingerprint modules, i.e., the two places where memory safety is actually decided. The ABI crate, exposing all 4,119 entrypoints, allows unsafe outright. The deny gates are real and better than nothing, but the membrane's most load-bearing modules are exempt from them by documented convention. A `SAFETY` comment is a promise, not a proof, and promises in the allocator are exactly where the CVE history lives.

**[HIGH] 30.7% of the surface is glibc with a hall monitor.** For 1,264 symbols the membrane validates inputs and then calls the host. The safety claim for nearly a third of the ABI is therefore bounded by glibc's behavior — the very behavior the project exists to improve on. This is honest (the matrix says so) and strategically necessary (L1 staging), but it means the headline "memory-safe libc" describes at most 69.3% of the surface, and the dangerous remainder is validated, not replaced.

**[HIGH] The velocity that built this can silently unbuild it.** Four commits authored 2026-06-26 (committed 2026-08-03, replayed onto the branch so `--since`/`--until` queries miss them) silently deleted shipped work: ~1,300 lines twice over, the entire f128 math engine plus the C23 `fromfp` family (−6,914), and the qsort fast/radix lanes. Three banked benchmark wins stood falsely in the ledger for ten weeks after their code was gone, retracted 2026-09-01. The project caught this itself and published a superb audit — which is admirable — but the defect class is structural: at ~25–50 commits/day with ~49% agent co-authorship, silent deletion is a background process, not an incident. The audit found nine deletion-excess candidates; adjudication is ongoing.

**[MEDIUM] The runtime math is asserted, not demonstrated.** 71 kernel files, 12 design docs, a mandated 44-family math stack with branch-diversity rules — and a kill-switch (`FRANKENLIBC_RUNTIME_MATH=off`) so operators can disable the whole thing. Existence is verified; production value is not: no published measurement isolates what the kernels buy (see Next Step 5), and the AGENTS.md math mandates read as process theater until the numbers land. The README's own FAQ — "is the runtime math real code or just naming theater?" — shows the maintainer knows exactly how this looks.

**[MEDIUM] Interposition's ceiling is physics, not effort.** `LD_PRELOAD` cannot see static binaries, setuid/setgid binaries, or anything that runs before the preload resolves. The DEPLOYMENT.md admits all of this plainly. It bounds every production claim permanently at L1 — which makes the L2/L3 roadmap load-bearing for any claim beyond "validated shim," and L2/L3 do not exist.

**[MEDIUM] CI is red at the pin and the smoke artifact is 3.5 months old.** Main CI run #8942 failed on its first step (the core `scripts/ci.sh` gate) with every substantive job skipped; the failure cause was not established from API-visible data. The curated smoke battery's checked-in summary is dated 2026-06-03. 588 gate scripts encode admirable contracts, but the two highest-visibility signals — CI status and the smoke artifact — are both stale or red at the assessment pin.

**[MEDIUM] Documentation lags velocity structurally.** The README's architecture diagram claims 554 shell scripts; the tree has 631 (588 of them `check_*.sh`). The README's "Current State" was reconciled 2026-09-18 — four days before the pin, an eternity at this commit rate. At ~50 commits/day, any prose claim about the repo decays within days; the machine-checked artifacts (matrix, ledger, JSON summaries) are the only claims that keep up.

**Bear-case steelman.** FrankenLibC is a one-maintainer, agent-assisted research prototype whose safety-critical modules are exempt from its own unsafe gates, whose benchmark wins were once silently deleted without the gates noticing, whose CI is red, whose license forbids evaluation by the labs that could validate it, and whose 30.7% host-backed surface means the "safe libc" still calls glibc a thousand times per process. The runtime-math program may be naming theater with a kill-switch. What survives this telling is the evidence machinery — the strongest this packet documents — and even that proved insufficient to catch its own silent deletions for ten weeks. An acquirer or adopter is buying process discipline and a 4,119-row classification, not a libc.

---
### 4.8 Maintenance and succession

**Bus factor: one.** The contributor API shows a single account (Dicklesworthstone) across 11,272 commits since the repo's creation on 2026-02-09. There is no CONTRIBUTING file, no second maintainer, no organizational home. Issue tracking is the bespoke in-repo bead system (`.beads/issues.jsonl`: 7,422 closed, 32 open), which means the project's working memory lives in a JSONL format that dies with the maintainer's tooling choices.

**Authorship is substantially agentic.** In the 2026-08-18→2026-09-22 window (895 commits sampled via blobless history), 437 commits (48.8%) carry `Co-Authored-By` trailers: 406 Claude Opus 5, 19 Claude, 9 Claude Opus 5 (1M context), 3 Grok. The pin itself has no trailer, but its message cites a candidate validation workflow run. Trailer counts measure commits touched, not lines written — but at ~25 commits/day in the window, the honest description is a human directing an agent fleet. The silent-deletion audit (§4.7) is the documented failure mode of exactly this arrangement: agents that add brilliantly and delete silently, with commit messages "accurate about what they *added*" while the deletions go unremarked.

**Succession assets and liabilities.** On the asset side: the bead tracker is extraordinary institutional memory (7,422 closed beads with evidence-closure discipline — the AGENTS.md rule that "a bead may not close on a gate whose compilation was not observed" exists because four beads did exactly that); the AGENTS.md file is a 2,109-line operating manual that would let a competent successor run the project; the machine-checked artifacts (support matrix, replacement levels, packaging spec, smoke summary) are self-describing. On the liability side: all of it is bespoke — bead IDs, gate scripts, claim taxonomies — with no community, no releases, no tags, and a license that forbids the largest AI labs from even analyzing the code. A successor inherits a magnificent private laboratory, not a project.

**License — read verbatim at the pin.** The header is `MIT License (with OpenAI/Anthropic Rider)`. The rider's restricted parties: "OpenAI, L.L.C. ('OpenAI'), Anthropic, PBC ('Anthropic'), their affiliates, and any person or entity acting directly or indirectly for the benefit of, or under the direction of, any of the foregoing." The restricted "Use" is defined to include "copying, modifying, merging, publishing, distributing, sublicensing, selling, transferring, making available, hosting, deploying, executing, **benchmarking, testing, analyzing, indexing**, or incorporating the Software or any Derivative Works into any dataset, training corpus, evaluation harness, or pipeline for machine learning or other automated systems" (emphasis added). The rider applies to the software and all derivative works, must remain unmodified on distribution, and breach "will automatically terminate all permissions." [Inference, not legal advice] This is not an OSI open-source license: it discriminates against named persons and against fields of endeavor (benchmarking, training, evaluation). GitHub classifies it as `Other / NOASSERTION`. Practical consequence: any organization touching this code must clear the rider, and the labs most capable of auditing it cannot touch it at all.

### 4.9 NODUS factsheet

| Dimension | Score | Rationale |
|---|---|---|
| Technology Readiness Level | **4** | Lab validation: real binaries run under the interposed library in a curated battery (June artifact), same-invocation benchmarks against live glibc, 66 fuzz targets. No production deployment, no release, no independent replication. The interpose artifact works in the lab; nothing has left it. *Moves to 5 on:* an independent rerun of the smoke battery at a fresh pin, or any production deployment. |
| Strategic relevance | 4/10 | Memory safety at the libc boundary is a real strategic problem; an interposition validator is a plausible wedge. But L1 caps the strategy, and the rider blocks the ecosystem that would make it strategic. *Moves up on:* a credible L2 path or a second institutional adopter. |
| Potential impact | 3/10 | If the claim machinery were adopted suite-wide or industry-wide, the impact would be methodological, not from this libc. The libc itself impacts nobody until L2 exists. *Moves up on:* extraction of the claim-field contract into a reused standard (see Next Steps 1–3). |
| Implementation feasibility | 2/10 | L2/L3 (standalone replacement) is a multi-year, multi-engineer undertaking; current resourcing is one human plus agents. The 30.7% host-backed tail is the long pole, and it is host-backed for good reasons (locales, NSS, dynamic linker interplay). *Moves up on:* a second maintainer or a funded team. |
| Time to mainstream | 2/10 | No credible path to production adoption within 24 months: no releases, red CI at pin, bus factor one, license friction. *Moves up on:* a tagged release with green CI. |
| Collaboration readiness | 2/10 | Superb documentation for a successor, zero community infrastructure, and a license that forbids the most capable potential collaborators from analyzing the code. *Moves up on:* rider narrowed or removed, plus a second committer. |

**NODUS ring: Explore.** Substantive enough to track closely (the claim machinery, the 69.3% native surface, the live benchmark discipline), unproven enough that no commitment is warranted. Revisit triggers in §4.11.

### 4.10 Wardley placement (assessor inference — placements are analytical judgments, not measured positions)

- **Commodity:** the C ABI itself, POSIX semantics, glibc behavior as the compatibility oracle. Nobody differentiates here; the project correctly treats glibc as the incumbent to be measured against, not beaten in the abstract. *Moves:* nowhere — commodities don't move.
- **Custom-built (product-ish):** the interposition layer + policy membrane + claim/evidence machinery. The membrane is the differentiator; the claim-field contract (symbol status ≠ semantic parity ≠ replacement level ≠ freshness) is approaching a product in its own right — it could be extracted and applied to any of the 44 repos tomorrow. *Moves toward product:* if a second project adopts the claim-field contract or the campaign-win taxonomy (see Next Steps 1–3).
- **Genesis:** a standalone Rust libc replacement (L2/L3) and the runtime-math control plane as a packaged discipline. The 44-family math mandate is genesis-stage process: possibly a genuine innovation in agent-directed verification, possibly theater. *Moves toward custom-built:* a published runtime-math isolation measurement (Next Step 5) or an L2 prototype; *moves toward irrelevance:* another year with neither.

### 4.11 Trajectory — 12 / 24 / 60 months [Inference]

**12 months.** Most likely: continued high-velocity interposition work, native share of the matrix climbing from 69.3% toward 75–80%, more campaign wins banked, the silent-deletion audit methods hardening into standard practice. CI presumably returns to green; a v0.1.0-style release would be the credibility event to watch for. The runtime-math program either produces an isolating measurement (Next Step 5) or remains asserted. [Inference]

**24 months.** The fork in the road: either L2 (standalone `.so`) materializes with a credible no-host-glibc boot path, or the project settles into being the best-instrumented L1 validator in the world — valuable as methodology, not as infrastructure. L2/L3 are not undesigned: PROPOSED_ARCHITECTURE.md already frames them as "distinct contract surfaces" with stated promotion requirements (replacement-closure evidence for symbols touching implementation-private layouts like `FILE` and loader state; errno, TLS finalization, and process-exit semantics declared per mode and per packaging level). What does not exist is an implementation, a prototype, or a schedule — the requirements are specified, the build is not. Bus factor remains the existential risk; a second maintainer or institutional home would change the trajectory more than any technical milestone. [Inference]

**60 months.** If L2/L3 ever land with the claim machinery intact, this becomes a genuine alternative-libc story and the interposition years read as the world's most careful on-ramp. More likely: the libc remains a research artifact while the evidence discipline (campaign-win taxonomy, deletion-excess audits, claim-field separation) gets adopted elsewhere — the methodology outlives the artifact. The rider, if unchanged, ensures the artifact's influence stays indirect. [Inference]

**Revisit triggers:** (1) main CI green at a pin with substantive gates executed; (2) a tagged release or published artifact; (3) L2 standalone `.so` demonstrated without host glibc; (4) an independent third-party benchmark or deployment; (5) a second maintainer with merge rights; (6) a published runtime-math isolation measurement.

### 4.12 Verdict and ring

**Verdict: TRL 4, NODUS Explore.** FrankenLibC is the most rigorously self-audited system this program has assessed [Inference, High] and one of the most honest measurement regimes in open-source systems work: it publishes its losses, retracts its stale wins, audits its own silent deletions, and separates every claim into machine-checked fields so that no statement outruns its evidence. It is also, at the pin, a red-CI L1 interposition prototype with a 3.5-month-old smoke artifact, a one-person bus factor, ~49% agent co-authorship, a license that forbids evaluation by the labs best equipped to perform it, and a 30.7% host-backed surface that bounds every safety claim.

**Cross-cutting lenses.** *Decoupling:* FrankenLibC advances the decoupling of validation from implementation — the `WrapsHostLibc` path proves you can validate a call without reimplementing it — and the decoupling of claims from code via machine-checked metadata. *Methodology-export:* if the libc fails, the claim-field contract, the campaign-win taxonomy, and the deletion-excess audit method survive as portable artifacts; they are the strongest candidates in the suite for extraction. *Asupersync:* verified dev/tooling-only, not a runtime dependency (§4.4, claim 13). *Rider:* quoted verbatim and assessed as strategy in §4.8 and deepening question 6. The methodology is the durable export; the libc is the demonstration vehicle. Explore — with the revisit triggers above as the promotion criteria.

### 4.13 Limitations and open questions

1. **No build or execution was performed.** All code claims rest on tree inspection and checked-in artifacts, not on compiling or running anything. The smoke battery, fuzz targets, and benchmarks were read, not reproduced.
2. **Shallow history.** The clone was blobless with a sampled window (895 commits, 2026-08-18→2026-09-22); lifetime commit counts come from the GitHub API. The silent-deletion audit's claims were read in AGENTS.md, not re-executed.
3. **CI internals not visible.** The pin's CI failure was observed via the Actions API (job names, conclusions, durations); the failing log's contents were not retrieved, so the root cause is unestablished.
4. **Mathematical validity not audited.** The runtime-math kernels' correctness and the "44 math families" doctrine were verified for existence and process, not for mathematical soundness — that audit would require domain expertise beyond this packet.
5. **Coverage search is recall-limited.** "No independent coverage" reflects web search results, not proof of absence; non-English or non-indexed discussion could exist.
6. **License analysis is not legal advice.** The rider's interpretation (non-OSI, discriminates against named parties and fields) is the assessor's inference from the verbatim text.
7. **Open:** what fraction of the 4,119 symbols' `Implemented` rows have differential conformance evidence vs. presence-only? The matrix classifies; the conformance depth per row was not sampled.
8. **Open:** does the 66-target fuzz corpus run in CI at the pin, and with what coverage? The targets exist; their execution status at the pin is unverified.

---

## Deepening Questions

**1. Provenance.** The repo records extensive provenance: every commit carries author and co-authorship trailers (437 of 895 sampled commits in the recent window carry AI co-authorship), the bead tracker ties each claim to evidence artifacts by bead ID, the smoke summary carries run ID, bead, schema version, and check timestamp, ledger entries carry dates and commit hashes, and the support matrix is machine-generated from the tree. What is missing for portability: there are no signed attestations of any kind — no signed tags (there are no tags at all), no Sigstore/cosign signatures, no SLSA provenance — and the bead JSONL is a bespoke format no external tool reads. Making attestation portable would mean emitting standard in-toto/SLSA attestations per checked artifact and signing releases; none of that exists yet, because there are no releases to sign.

**2. The embeddable unit.** The smallest useful piece is the claim-field contract plus the campaign-win/self-speedup taxonomy: a documentation-and-script convention (the support-matrix schema, the check-script gate pattern, the rule that a win needs the incumbent live in the same invocation) adoptable by any project at the cost of writing per-claim metadata and a CI gate — days of work, no Rust required. The next smallest is the membrane's per-call validation policy as a standalone `LD_PRELOAD` shim for a *subset* of dangerous functions (say, just the string and memory functions), adoptable without the 4,119-symbol surface at the cost of maintaining the interposition glue and accepting the framing overhead the project's own profiles put near 30% of process self-time. The full libc is not embeddable — it is the whole repo, and it still needs host glibc underneath.

**3. Unexercised option value.** The architecture holds several capabilities it has not used: the 66-target fuzz corpus exists but its execution status at the pin is unverified — if it is not running in CI, that is dormant assurance; the Gentoo Portage-hook lane is the only deployment story with per-package operational granularity (allowlists, blocklists, per-phase control) and has no known external user; the structured runtime log (`FRANKENLIBC_LOG` JSONL) is a per-process call trace that could feed anomaly detection but is currently just a log; and the `WrapsHostLibc` membrane path is a pure-observability product (validate and log without replacing — the L0 value proposition) that was never productized. What unlocks each: CI execution of the fuzzers, one external user of the Portage lane, and a documented observability use case with a before/after incident story.

**4. Benchmark honesty.** The campaign wins (`strftime` 0.5406×, `wcsftime` 0.1113×) have the best-documented methodology in the packet — dlmopen-protected incumbent, A/A null controls, bootstrap CIs, 200,000-case differentials — and would most likely survive an independent rerun; but they are microbenchmarks on stateless formatting functions, load-bearing for nothing except the claim that the methodology works. The numbers load-bearing for the thesis do not exist yet: whole-workload membrane overhead (the ~30% framing figure is a profile attribution, not a workload benchmark) and the native-share trend that justifies L2. The published losses (`malloc_free` 5.644× slower) would survive trivially — losses are cheap to reproduce. And the honest caveat: nothing here has been independently rerun, so "would survive" is itself an inference from methodology quality, not an observation.

**5. The governance path.** The credible route from one maintainer to an institution runs through the succession kit that already exists: the 2,109-line AGENTS.md operating manual, the bead tracker with evidence-closure discipline, and the 588 machine-checked gate scripts mean a second maintainer could run the project from the docs — which is genuinely rare. The institutional route is a security-focused company or lab adopting the L1 validator for hardened fleets, or a grant funding the methodology-export work (Next Steps 1–3). What breaks first if velocity decays: the prose docs (already four days stale at the pin), then the ledger's freshness — stale wins stood ten weeks *at full velocity*, so at low velocity the claim/evidence gap widens silently — then the nightly toolchain pin rots. The machine-checked artifacts break last, which is exactly why they are the durable export.

**6. The license as strategy.** The rider excludes OpenAI, Anthropic, their affiliates, and anyone acting directly or indirectly for their benefit or under their direction, from an enumerated "Use" that explicitly includes benchmarking, testing, analyzing, indexing, and incorporating the software into training corpora or evaluation harnesses — with breach auto-terminating all permissions. As strategy it serves exactly one plausible goal: keeping the 4,119-symbol surface and its test corpus out of the largest labs' training data and benchmark pipelines. It sabotages everything else: independent evaluation by the best-equipped auditors, enterprise adoption (no legal department clears named-party discrimination plus auto-termination), and any acquisition. And it is internally incoherent — roughly half of recent commits are co-authored by the excluded party's models. A narrower instrument, such as a training-data opt-out, would preserve the defense without the sabotage; the current rider reads as strategy written against phantoms at the expense of real counterparties. [Legal interpretation is the assessor's inference, not legal advice.]

**7. Agent-era fit.** The concrete agent workload that would pick this over glibc is sandboxing untrusted AI-generated code: an `LD_PRELOAD` membrane that validates libc calls, denies dangerous patterns in hardened mode, and emits structured JSONL traces is a natural complement to container sandboxes for agent-executed code. What has to become true first: the overhead story (agents are latency-sensitive and the framing tax needs a workload-level number, not a profile attribution), the coverage story (the sandbox's riskiest surface is exactly the 30.7% that calls through to glibc — validation without replacement is half a sandbox), and the trust story (the watcher must be more trustworthy than the watched, and the unsafe exemptions in the allocator and fingerprint modules cut against that). The irony the packet keeps returning to is also the proof of concept: the maintainer's own agent fleet already works inside this codebase every day.

**8. The kill test.** The core thesis is "validate the boundary you cannot replace": per-call interposition validation is a viable path to a safer userspace without replacing glibc. The falsifying experiment is a whole-workload benchmark showing the membrane's framing overhead makes realistic programs unacceptably slow *and* that the overhead is structural rather than optimizable — if validation costs more than the safety is worth, L1 has no adopters and L2 inherits the same membrane. The competitor-move version: musl or glibc shipping a first-party validated-call layer, or the kernel exposing a cheap call-validation mechanism, would strand the interposition approach overnight. The maintainer already thinks in falsification terms — the runtime-math kill-switch (`FRANKENLIBC_RUNTIME_MATH=off`) is a built-in kill test for the strangest subsystem — but the thesis-level kill test, the workload benchmark with the membrane on versus off, has not been run.

---

*Packet version: v7 · Evidence labels inline · All URLs verified resolving 2026-09-22 at the pinned commit · Supersedes the 2026-09-21 draft in every quantitative claim.*
