Agent Work Hub · New task

#17 P5 — Targeted invalidation-model correction: the env key, and a healthy (B) baseline

project
hyperhash
status
cancelled
holder
created by
owner
runner kind
laptop
lease
budget usd
10.0
created
updated

Instructions

Return rule "P5 detects stale-evidence risk" (`docs/architecture/ARCHITECT_RULING_2026_10_04_OVERNIGHT_TASKS_TAKEN.md` §1): a targeted invalidation-model correction; no broad implementation. Surface: `verification_turnaround`.
Read first: `tools/test_verification_turnaround_findings.md` (§4 the env-key gap, §6 the recommendation),
`docs/architecture/VERIFICATION_TURNAROUND_PREREGISTRATION_2026_10_04.md` (FROZEN; do not edit), and the harness
`tools/test_verification_turnaround.py` + `tools/p5_mutations.py`.
1. **Correct the env key.** The §3.1(4) key misses 173 env var names the entries' dependency sources read. Change the
   harness's key to cover them: prefer the traced set of env names each entry actually reads; if tracing cannot see
   every read, fall back to the whole `ARC_*` / `SUDOKU_*` / `GCOS_*` / `SCAN_*` / `CLUTRR_*` / `MATH_LAB_*` families.
   Say which you chose and why. This is a harness change only, not a cache system.
2. **Re-validate on the same 7 entries and the same mutation set**, plus env mutations that use a value other than `1`
   (e.g. `0` for hatches, which are bare truthiness) and at least one non-literal read if any exists. Every HIT is
   re-run; report false accepts, stale hits, missed invalidations and over-invalidations.
3. **A healthy (B) baseline.** The dataset files missing from this checkout were copied in on 2026-10-04 (ARC-AGI,
   Sudoku, math). Re-run the full suite once (`.venv/bin/python -m pytest -n 4`) and record wall clock, load average
   and failures; read failures from junit XML, and check them against `tools/close_gate_baseline.json` before calling
   any of them new. If (B) still cannot be established, report the missing instrumentation. Do not edit the baseline.
   Note: `close_gate` (B) was reported failing on an Inv-17 row from
   `tools/test_experience_to_structure_probe_results.json`; report whether it reproduces, do not fix it here.
Stop on any false accept or unrecognized invalidation case. Do not weaken any verification requirement. A cached
result never replaces close or release evidence.
Add a dated section to `tools/test_verification_turnaround_findings.md` (MEASURED and PROJECTED kept separate) and a
one-line recommendation: proceed to an implementation proposal limited to the validated boundaries, or not.
Submission: the recommendation, the counts, the (B) wall clock, the commit, and for any failure its class (protocol
issue, infrastructure issue, insufficient evidence, or actual hypothesis failure). Claim with
`tools/loop_budget.py claim --surface verification_turnaround --experiment p5_env_key_correction --lease-min 240`,
release when banked, commit, push.

Submissions (0)

None yet.