Agent Work Hub · New task

#15 P5 — Verification baseline and cache correctness (bounded validation)

project
hyperhash
status
done
holder
laptop-runner
created by
owner
runner kind
laptop
lease
budget usd
5.0
created
updated

Instructions

Medium priority: measurement and a small test. Surface: `verification_turnaround`. It does not need the census to finish.
Read first: `docs/architecture/ARCHITECT_RULING_2026_10_04_OVERNIGHT_TASKS_TAKEN.md` (§2 is this task's full text, the
architect's "Task 3"; §1 has the return rules and failure classes), `docs/architecture/VERIFICATION_TURNAROUND_PREREGISTRATION_2026_10_04.md`
(FROZEN; do not edit), and CLAUDE.md. If another hub task is running on this machine, work in your own git worktree.
Objective: a trustworthy verification-turnaround baseline, and a test of whether the proposed layered-evidence cache
correctly invalidates stale evidence.
- Inspect the current verification suites and the preregistration's cache dependency model.
- Pick a representative, bounded set of verification runs. Record actual execution times, cache hits, invalidations
  and total turnaround.
- Timing hygiene: the ARC-1 census may be running on this machine. Record the load average with every timing and say
  which timings overlapped a census run; prefer timing when it is not running.
- Inject a small, controlled set of changes that should invalidate cached evidence: relevant code, tests,
  configuration (including `RESULT_AFFECTING_ENV_VARS`) and dependencies. Test that none of them lets stale evidence be
  accepted.
- `close_gate`'s set containment and the four honesty gates are untouched. A cached result never replaces full close
  or release evidence. Build no broad caching system: only the harness this test needs.
Stop conditions: stop on any false accept or unrecognized invalidation case and report it. Do not weaken any
verification requirement to improve performance. If the baseline cannot be established reliably (load, missing
timings), report the missing instrumentation rather than inventing a comparison.
Deliverables, in `tools/test_verification_turnaround_findings.md` (headline table first, ≤ ~200 lines): baseline
measurements; invalidation test results; false-accept and stale-evidence findings; a recommendation on whether to
proceed with implementation; a rough estimate of time savings with MEASURED and PROJECTED kept in separate columns.
Submission: the recommendation, false accepts and stale hits (counts), the commit, and for any failure its class
(protocol issue, infrastructure issue, insufficient evidence, or actual hypothesis failure). Use `gtimeout` on macOS.
Claim with `tools/loop_budget.py claim --surface verification_turnaround --experiment p5_cache_validation --lease-min
240`, release when banked, commit, push.

Submissions (1)

laptop-runner
session exited with code 1 without submitting (timeout: no)

```
Permission allow rule (.claude/settings.json): Write(**) is not matched by file permission checks — only Edit(path) rules are. Use Edit(**) instead (Edit rules cover all file-editing tools).
Permission allow rule (../hyperhash/.claude/settings.local.json): Bash(mv SESSION_SUMMARY*.md docs/sessions/) has a wildcard before the rest of the command, so it also matches any options inserted at that position and approves them without a prompt. Replace that * with the exact value you mean, or only use * after the subcommand.
Permission allow rule (../hyperhash/.claude/settings.local.json): Bash(mv MASTER_STATUS_*.md DEPLOYMENT_SUMMARY.md PROJECT_STATUS.md docs/status/) has a wildcard before the rest of the command, so it also matches any options inserted at that position and approves them without a prompt. Replace that * with the exact value you mean, or only use * after the subcommand.
Error: Exceeded USD budget (5)
```