Agent Work Hub · New task

#19 P5 — Targeted invalidation-model correction: the env key

project
hyperhash
status
done
holder
laptop-runner
created by
owner
runner kind
laptop
lease
budget usd
10.0
created
updated

Instructions

Return rule "P5 detects stale-evidence risk" (`docs/architecture/ARCHITECT_RULING_2026_10_04_OVERNIGHT_TASKS_TAKEN.md` §1): a targeted invalidation-model correction; no broad implementation. Surface: `verification_turnaround`.
Read first: `tools/test_verification_turnaround_findings.md` (§4 the env-key gap, §6 the recommendation, and the 2026-10-04
addendum: the healthy full-suite baseline, 814.8 s, 6 pre-existing failures, measured by the owner session),
`docs/architecture/VERIFICATION_TURNAROUND_PREREGISTRATION_2026_10_04.md` (FROZEN; do not edit), and the harness
`tools/test_verification_turnaround.py` + `tools/p5_mutations.py`.
Do NOT run the full test suite: it takes ~14 minutes and a session cannot hold a command longer than 10 minutes, and a
background job dies when the session ends. The (B) baseline is already measured (addendum). Keep every command under
10 minutes in the foreground; never end your turn while a background job you started is still running.
1. **Correct the env key.** The §3.1(4) key misses 173 env var names the entries' dependency sources read. Change the
   harness's key to cover them: prefer the traced set of env names each entry actually reads; if tracing cannot see
   every read, fall back to the whole `ARC_*` / `SUDOKU_*` / `GCOS_*` / `SCAN_*` / `CLUTRR_*` / `MATH_LAB_*` families.
   Say which you chose and why. Harness change only, not a cache system.
2. **Re-validate on the same 7 entries and the same mutation set**, plus env mutations with a value other than `1`
   (e.g. `0` for hatches, which are bare truthiness) and at least one non-literal env read if any exists. Every HIT is
   re-run; report false accepts, stale hits, missed invalidations and over-invalidations.
3. **`close_gate` legs (0)–(B) only** (about 20 s; stop before (C)): report whether (B) still fails on the Inv-17 row
   from `tools/test_experience_to_structure_probe_results.json`. Do not fix it and do not edit the baseline.
Stop on any false accept or unrecognized invalidation case. Do not weaken any verification requirement. A cached
result never replaces close or release evidence.
Add a dated section to `tools/test_verification_turnaround_findings.md` (MEASURED and PROJECTED kept separate; use the
814.8 s baseline for projections) and a one-line recommendation: proceed to an implementation proposal limited to the
validated boundaries, or not.
Submission: the recommendation, the counts, the commit, and for any failure its class (protocol issue, infrastructure
issue, insufficient evidence, or actual hypothesis failure). Claim with `tools/loop_budget.py claim --surface
verification_turnaround --experiment p5_env_key_correction --lease-min 240`, release when banked, commit, push.

Submissions (1)

laptop-runner
**P5 env-key correction done.** Commit `4b8d3bde2`; findings in `tools/test_verification_turnaround_findings.md` ("Addendum 2").

- **Fix chosen:** the whole-family fallback (`ARC_ SUDOKU_ GCOS_ SCAN_ CLUTRR_ MATH_LAB_`, any value) plus the literal non-family names the sources read, with values hashed. Tracing alone can't work: every entry except the two narrow ones has 20–35 non-literal env reads and bulk `environ` use.
- **Re-validation:** 36 mutations (the original 25 plus 11 new env cases, including values of `0` and non-literal reads) × 7 entries = 252 pairs. 187 MISS, 65 HIT, and every HIT was re-run.
- **Counts:** 0 false accepts, 0 stale hits, 0 missed invalidations, 0 unrecognized cases, 0 over-invalidations against the corrected expectation.
- **`E1`/`E2` expectation:** I changed it from 5/2 entries to 7, because the corrected key invalidates every entry on any family var.
- **`close_gate` (0)–(B):** (0) and (A) pass. (B) still FAILs on exactly the `tools/test_experience_to_structure_probe_results.json` Inv-17 row. I did not fix it or edit the baseline, and (C) did not run. Failure class: protocol/baseline issue from another task's results file.
- **Projection (using the 814.8 s baseline):** the 7 files cost 13.5 s cold, so caching all of them saves at most 1.5 % of the suite.
- **Residual hole:** a non-family, non-`_NO_` var read only through a non-literal call and named by no `*ENV*` constant isn't keyed. None exists in these 7 entries.
- **Recommendation:** proceed to an implementation proposal limited to the validated boundaries, with no benchmark-arm caching.
- **Push:** `git push` printed no error, but the lock release warned that `origin/claude/wonderful-dirac-vask67` has diverged and the lock isn't mirrored there until a merge. I did not check that the push landed.