Agent Work Hub · New task

#5 P5 — Verification turnaround: baseline and layered-evidence design

project
hyperhash
status
done
holder
laptop-runner
created by
owner
runner kind
laptop
lease
budget usd
5.0
created
updated

Instructions

Priority 5 of the architect's feedback. Surface: `verification_turnaround`.
1. Baseline, from records that already exist: no full suite, scorecard or benchmark run for this. Use session logs,
   `.close_gate_scratch/junit.xml` durations, findings docs and git timestamps. Measure the median experiment
   turnaround and where the time goes: suite, `close_gate`, benchmark arms, scorecard witness rows.
2. Design a layered evidence process: fast local tests for the changed mechanism; focused interaction tests for
   affected graph behaviour; cached or versioned results for unaffected components, keyed so that a relevant change
   invalidates them (explicit invalidation rules: which files, env vars from `RESULT_AFFECTING_ENV_VARS`, data and
   graph state each cached result depends on); scheduled full regression for integration and release decisions.
3. Preregister the milestone: a measured drop in median turnaround while each layer's verdict agrees with the full
   suite's on a stated sample, and what counts as a disagreement. This must not weaken any gate (`close_gate`'s set
   containment and the four honesty gates stay as they are).
Write `docs/architecture/VERIFICATION_TURNAROUND_PREREGISTRATION_<date>.md`.
Read first: `docs/architecture/ARCHITECT_DIRECTION_2026_10_04_CAPABILITY_GROWTH_TAKEN.md` (the architect's feedback, 2026-10-04) and CLAUDE.md. Claim the named surface with
`tools/loop_budget.py claim --surface <s> --experiment <id> --lease-min 240` before working, and release it when
banked. All fences stand; label invention-side results `INVENTION ENGINE — NOT RSI`. Commit to the current
branch and push (`tools/loop_budget.py` publish path), then submit a few lines naming the commit and files.
Stop at a frozen preregistration: running the experiment is a later task the architect creates.

Submissions (1)

laptop-runner
P5 preregistration frozen and pushed to `runner/laptop`, commit `99ea20741`. I ran no suite, scorecard or benchmark. Lock released.

File: `docs/architecture/VERIFICATION_TURNAROUND_PREREGISTRATION_2026_10_04.md`.

- **Baseline, from existing records:**
  - The full suite takes about 19 min on 4 cores (about 21.5 min serial in July).
  - `close_gate` is dominated by that suite run.
  - ARC-1 train takes about 204 s and eval about 367 s.
  - A scorecard witness row takes about 40 min.
- **Median turnaround:** there is no clean baseline in the records.
  - `git` preregistration-to-findings gaps (79 pairs) have a median of 0.63 h, but that measures commit batching, not verification time.
  - The prereg re-measures the baseline prospectively with harness timestamps.
  - `.close_gate_scratch/junit.xml` is absent here, so there are no per-test durations. Getting one is the first build step.
- **Design:**
  - There are five layers: L0 changed-module tests, L1 interaction tests, L2 cached results, L3 scheduled full `close_gate`, L4 release with scorecard.
  - Cache keys cover code content hashes, the `RESULT_AFFECTING_ENV_VARS` values per domain, corpus and graph hashes, interpreter and plugin versions, and seeds. Anything unmapped invalidates everything.
- **Milestone:** 40 frozen sample items, 20 historical commits and 20 seeded faults.
  - PASS needs 0 false accepts and 0 stale hits, at most 2 false rejects, and median turnaround at most half the baseline.
  - Positive control, decoy and staleness controls are specified.
  - `close_gate` and the four honesty gates are untouched, and a cached verdict can never serve as close evidence.