Agent Work Hub · New task

#3 P3 — Failure-to-improvement loop: design and preregistration

project
hyperhash
status
done
holder
laptop-runner
created by
owner
runner kind
laptop
lease
budget usd
8.0
created
updated

Instructions

Priority 3 of the architect's feedback, the biggest strategic change. Surface: `failure_growth_loop`.
Write the design and a preregistration (CLAUDE.md §7: commit the prereg before any runner exists) for a loop that turns
failed tasks into proposed, tested, retained and reused graph-level improvements.
- For each failure the loop records, as graph-resident structure (not a one-off diagnosis script): the missing
  distinction or transformation; whether the input carried the information; whether the graph could express it;
  whether a right candidate was generated but ranked low; the smallest graph-level change that could fix it; and
  whether that change helps held-out tasks without losing existing ones.
- Pick one manageable family with many cheap, independent failures. Say why its members are independent and where
  the failures come from (no manufactured populations).
- Proposals must come from the system's own search over failure evidence, never from an abstraction the experiment
  author supplies (`K-1` / `B-iii`, `NO-MANUFACTURED-ORIGINATION`). State the derivation operator.
- The pass criterion is the cycle's test: at least one useful, previously absent graph abstraction made from failure
  evidence, validated on held-out tasks and reused, with no human specifying it. A fail must name which stage of the
  loop is missing. Include a positive control and a decoy.
- Say which parts of the loop are hand-engineered now and how each moves onto graph-resident, reusable structure.
Write `docs/architecture/FAILURE_GROWTH_LOOP_PREREGISTRATION_<date>.md`.
Read first: `docs/architecture/ARCHITECT_DIRECTION_2026_10_04_CAPABILITY_GROWTH_TAKEN.md` (the architect's feedback, 2026-10-04) and CLAUDE.md. Claim the named surface with
`tools/loop_budget.py claim --surface <s> --experiment <id> --lease-min 240` before working, and release it when
banked. All fences stand; label invention-side results `INVENTION ENGINE — NOT RSI`. Commit to the current
branch and push (`tools/loop_budget.py` publish path), then submit a few lines naming the commit and files.
Stop at a frozen preregistration: running the experiment is a later task the architect creates.

Submissions (1)

laptop-runner
P3 preregistration committed at `3f5abae04` on `runner/laptop`, pushed to origin. The surface lock is released. No runner exists and nothing was run.

File: `docs/architecture/FAILURE_GROWTH_LOOP_PREREGISTRATION_2026_10_04.md`

- **Family:** the integer lab's super-linear tasks under `M-Σ`. It uses the 42 of 105 tasks unsolved in banked run `exp_244ef8685d60e60f`. Held-out tasks come from fresh seeds 31–40, and all inference is counted per generator block.
- **Failure records:** six mechanised predicates stored as properties and edges on `MEMORY` episode nodes in the lab learner's own graph file. No new node or edge types.
- **Derivation operator:** a new failure-side search, "residual-anchored reification". It takes partial fits, computes their residuals, and runs a skeleton cover across failures. It is `D-2` re-enablement, `INVENTION ENGINE — NOT RSI`, with no author-supplied abstraction.
- **Pass and fail:** pass needs P1–P5, including held-out gain with 0 lost and reuse by another generator's task. A fail names the missing stage: NO-RECORD, NOT-EXPRESSIBLE, NO-PROPOSAL, PROPOSALS-ARE-COINCIDENCE, NOT-RETAINED or NOT-REUSED.
- **Controls:** a reconstruction control that deletes a reused sum-entry and requires the loop to re-derive it, plus a swap decoy and a matched-random decoy.
- **Hand-engineered now:** a table maps each part to its graph-resident target and the evidence that it moved. Nothing is claimed as moved.
- **Cost:** capped at 4 CPU-hours behind a priced shakedown.

Two things to check:
- The push of `runner/laptop` succeeded, but `loop_budget.py release` warned that `origin/claude/wonderful-dirac-vask67` has diverged and was not mirrored. I did not merge or force-push it.
- I did not add a session-log line. `next-id` offered 100, which the cloud session's lock name suggests is already taken.