Agent Work Hub · New task

#6 P2 — Substrate after A-25: why near-inert on natural problems, and what would change it

project
hyperhash
status
done
holder
laptop-runner
created by
owner
runner kind
laptop
lease
budget usd
3.0
created
updated

Instructions

Priority 2 of the architect's feedback is already done: `A-25` ran and Stage A came out NULL
(`tools/test_math_lab_natural_findings.md`). The architect's open row is: "a reason to expect the natural-problem test to
overcome the current near-inert behavior". Surface: `m2_a25_oeis` (claim it; it was released when M2 closed).
Write a short note (≤ 150 lines) for the architect, `docs/architecture/A25_SUBSTRATE_DIAGNOSIS_<date>.md`, from banked
records only (no new runs). Cover where the substrate failed on OEIS (reach, form, ranking), which of those the P3
failure loop could act on, and whether Stage B is worth running. If the answer is that nothing would change it, say so.
No preregistration here: this is a reading of banked evidence.
Read first: `docs/architecture/ARCHITECT_DIRECTION_2026_10_04_CAPABILITY_GROWTH_TAKEN.md` (the architect's feedback, 2026-10-04) and CLAUDE.md. Claim the named surface with
`tools/loop_budget.py claim --surface <s> --experiment <id> --lease-min 240` before working, and release it when
banked. All fences stand; label invention-side results `INVENTION ENGINE — NOT RSI`. Commit to the current
branch and push (`tools/loop_budget.py` publish path), then submit a few lines naming the commit and files.

Submissions (1)

laptop-runner
**P2 done.** Commit `1fe02a538` (release `ef5487713`), pushed on `runner/laptop`. Note: `docs/architecture/A25_SUBSTRATE_DIAGNOSIS_2026_10_04.md` (127 lines). It is a reading of banked evidence only, with no new runs.

- **Finding:** Nothing banked gives a reason to expect Stage B to overcome the near-inert result, and I recommend not running it as registered. Stage A never exercised the substrate. Its precondition (an affine image of a shared core that `B0` declines) is too scarce: about 2.6 tasks per 600 against a bar of 10.
- **Reach is the binding failure.** `B0` is right on about 1.2 % of in-range OEIS sequences.
- **Ranking is not the failure.** 157 of the 183 wrong answers are sequences that fit a polynomial on all 8 train terms and then change form. That is an evidence-window problem.
- **P3 failure loop:** It would classify most OEIS failures as vocabulary demand. It cannot fix reach, and it cannot see the changed-form failures, because it reads train pairs only.
- **What would change it:** a base-language reach change (an ask, e.g. a recurrence primitive) and Milestone 4's evidence acquisition. I propose a read-only recurrence census as an ask and did not run it.
- **Caveats:** Counts come from the findings file and were not re-derived. The findings file is on the cloud branch (`ada8df11e`) and is cited by commit, because it is not on this branch yet.

I claimed and released `m2_a25_oeis`. `loop_budget.py` warned that `origin/claude/wonderful-dirac-vask67` has diverged, so the lock may not be visible to that branch until it merges.