Files
automaton/tasks/fix-context-sizing/VERDICT.md
T
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

37 lines
2.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Verdict: fix-context-sizing
## Status: PASS
**Completion Date**: 2026-06-22
## Summary
Tier 1 context-sizing layer landed: `vram_detect.py`'s `recommend_context` now applies headroom exactly once (was 3×); fake 8k/6k fallbacks and `max(0,...)` lying clamp removed; new `--loop-mode` CLI flag refuses unknown models and sub-16k available context per D13; JSON output exposes `available_context_kb` + `loop_mode_eligible` for the upcoming loop runner. `config.md` documents `## Loop Role Models`; `decompose.md` gains a 4k tier and an explicit `≤ 16k: REFUSE` floor.
## Phase Outcomes
- Research: SPEC.md produced and approved by user before implementation.
- Implement: code edits + IMPLEMENTATION.md produced; 15 new tests in `tests/test_context_sizing.py`; one pre-existing test (`test_recommend_context_api_model`) updated to assert the corrected single-headroom contract.
- Code Review: PASS — full spec conformance R1–R6 verified; 264/264 tests green.
- Bug Find: NO_BUGS_FOUND — 4 adversarial vectors probed, no defects.
- Adversarial Bug Find: NO_NEW_DEFECTS — 5 attacks from design's §6 model all produce correct refuse/warn behavior.
- Doc Review: PASS — `config.md` + `decompose.md` updated; doc cross-references verified.
- Referee: PASS — all phase artifacts present and consistent.
## Test Results
- New: 15 passed / 15
- Pre-existing updated: 1 (`test_recommend_context_api_model` — was asserting the bug; now asserts the fix)
- Full suite: 264 passed
- Self-consistency suite: 74 passed (no prompt path regressions)
## Findings
- All six SPEC requirements (R1–R6) implemented and guard-tested.
- Pre-existing test that encoded the old buggy behavior was properly updated; reason documented inline in commit message.
## Tasks for Review / Tie-Breaks
- None
## Remaining Issues (out of scope, tracked)
- Loop runner must `max(0, available_context_kb)` before scheduling — task 3 (`add-loop-runner`).
- `README.md` human-facing CLI doc for `--loop-mode` — Tier 2 cleanup.
- v1.1 dashboard "Loops" panel will surface `loop_mode_eligible` — v1.1.
## Score
+10