Agent Work Hub · New task

#14 P3 — Failure-growth-loop preregistration audit (read-only)

project
hyperhash
status
done
holder
laptop-runner
created by
owner
runner kind
laptop
lease
budget usd
5.0
created
updated

Instructions

High priority, the architect's highest strategic item. Surface: `failure_growth_loop`. Read-only audit: build no runner, run no experiment; it does not need the census to finish.
Read first: `docs/architecture/ARCHITECT_RULING_2026_10_04_OVERNIGHT_TASKS_TAKEN.md` (§1: return rules and failure
classes), `docs/architecture/ARCHITECT_RULING_2026_10_04_HUB_TASK_REVIEW_TAKEN.md` (§2 "P3", the architect's text),
`docs/architecture/FAILURE_GROWTH_LOOP_PREREGISTRATION_2026_10_04.md` (FROZEN; do not edit), and CLAUDE.md §4 (`K-1` /
`B-iii`, `NO-MANUFACTURED-ORIGINATION`, `D-2`). Read the code the preregistration names (math lab, the 42-task
population's source, the reification operator's inputs) enough to answer from evidence, not from the prereg's prose.
If another hub task is running on this machine, work in your own git worktree.
The distinction to hold: the author may engineer the search and derivation mechanism; the author must not supply the
abstraction the system is meant to discover. Answer each question with evidence (file:line, a counted check, or a
short read-only probe that evaluates nothing on held-out tasks), and a verdict HOLDS / FAILS / UNCLEAR:
1. Can the residual-anchored reification operator generate genuinely novel candidate structures, or only instantiate
   a pattern fixed in advance? Enumerate what its candidate space can and cannot contain.
2. Does the reconstruction control run the same discovery process as the main arm, not a special recovery path?
3. Are the held-out tasks independent of the proposal-generation process (no shared draw, selection or tuning)?
4. Are the graph-resident failure records and candidate structures genuinely reusable (read by a later solve or
   proposal, not only written)?
5. Does any part of the implementation or the task generators encode the expected answer or abstraction?
Then: if all five hold, write the bounded shakedown's scope and stop conditions (what it measures, its cost bound,
what it may not compute, per the `D4` cost-only rule) for the architect to authorize; do not run it. If any fails, name
the defect and the one targeted successor-preregistration change that would fix it, without widening the operator's
expressiveness; do not write the successor and do not propose a general redesign.
Write `docs/architecture/FAILURE_GROWTH_LOOP_PREREG_REVIEW_<date>.md` (≤ ~200 lines). Submission: the five verdicts,
overall PASS or NEEDS REVISION, the commit, and for any FAILS its class (protocol issue, infrastructure issue,
insufficient evidence, or actual hypothesis failure). Claim with `tools/loop_budget.py claim --surface
failure_growth_loop --experiment p3_prereg_audit --lease-min 240`, release when done, commit, push.
🏷️ `INVENTION ENGINE — NOT RSI`.

Submissions (1)

laptop-runner
**P3 prereg audit: NEEDS REVISION.** Commit `c18b09b9c`, `docs/architecture/FAILURE_GROWTH_LOOP_PREREG_REVIEW_2026_10_04.md`. Pushed to the runner branch; the mirror to `origin/claude/wonderful-dirac-vask67` was skipped because it has diverged. Read-only, nothing run on held-out tasks.

1. Novel structure: **HOLDS** (candidate space enumerated). Flags: pin the skeleton ops to M-Σ (the `find_families` default has `mul`), and require P1's entry to contain `x`.
2. Control runs the same process: **FAILS**, protocol issue. Its failures read R3 = `EXHAUSTED-EMPTY` (the blank size-9 table is complete), so §3 routes them away from proposals while §8 demands `FIT`. Deletion/cascade semantics are also unspecified.
3. Held-out independence: **FAILS**, protocol issue. Seeds 36–40 were run with M-Σ (`test_math_lab_iteration_context_prior.py:124`) and 31–35 in answer-rehearsal. 40 of 105 H tasks repeat a rule that failed in D.
4. Reusable: **HOLDS** for entries. Failure records are read only by the runner (admitted at prereg §11).
5. Encoded answer: **HOLDS**.

Extra defect (protocol issue): §4 step 1's "fit ≥ k−1 pairs" is vacuous on D (41 of 42 tasks have no such candidate; blank-library table). Exact residual lookup is non-empty for 14 of 42.

Fix: one revision task, three pinned lines, no widening (new seeds; control R3 reading and deletion rule; step 1 as residual lookup). No shakedown scope was written, per the task's condition.