Complete tasks 3-7: harden verdict parsing, outputs retention, base branch, linux schedule parity, claim loop task
CI / build (push) Has been cancelled

This commit is contained in:
Lap Tran
2026-06-24 10:31:49 -04:00
parent dd2726c0dd
commit e13513faaa
193 changed files with 14934 additions and 98 deletions
@@ -0,0 +1,45 @@
# Doc Review: harden-parse-verdict
## Docs touched
- `design/loops/technical.md` §7 — tick-flow step 7 (parse verdict):
added 4-line inline block documenting the defensive coercion (pass
string acceptance; score clamp + NaN/inf/non-numeric → 0.5).
- `design/loops/functional.md` §10 — Verifier Contract: annotated
`pass` (bool) definitive + runner accepts `"true"`/`"false"` strings;
annotated `score` (0.0–1.0) clamp + NaN/inf/non-numeric → 0.5 neutral.
- `CHANGELOG.md` — new `[unreleased]` "Fixed — `parse_verdict`
defensive coercion" block above the existing `add-state-loop-lock` and
`fix-harness-command-template` blocks.
## Docs NOT touched (intentional)
- `AGENTS.md`: parse_verdict is not a user-visible CLI surface; the
hardening doesn't change phase enforcement, `.state.loop`, or any
contract that harness integrators need to know. The Verifier Contract
lives in `design/loops/functional.md` §10; AGENTS.md already points to
design docs at the top. No edit.
- `README.md`: user-facing README doesn't enumerate `parse_verdict`
internals; loop monitoring table mentions verdicts as a concept, not
the parser. No edit.
- `prompts/loop-verifier.md`: contract was already `bool pass` + `score
0.0–1.0`. The hardening is belt-and-suspenders against malformed
output, not a contract change. The prompt's strict-JSON directive
stays authoritative. No edit.
- `templates/loops/self-improvement/loop.json`: no schema change. No
edit.
## Cross-references
- `tasks/add-loop-runner/BUG_REPORT.md` O6 — the original finding —
now closed by this task. The CHANGELOG entry explicitly references it.
- `tasks/add-loop-runner/ADVERSARIAL_BUG_REPORT.md` A6 — the score-
clamping observation — also closed by this task. The CHANGELOG entry
references the clamping.
- `tasks/harden-parse-verdict/BUG_REPORT.md` O3 — notes that clamping
improves plateau detection (a tighter `score_history` range makes
plateau more honest). Cross-referenced from the CHANGELOG.
## Verdict
Docs are in sync with the implementation. Proceed to referee.