Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

2.2 KiB

DOC_REVIEW: add-loop-templates-onboarding

Reviewed Documentation

  1. README.md -- new "Loop Engineering" onboarding section (Quick Start, Tick Cycle, Configuration, Monitoring, Halt/Resume)
  2. CHANGELOG.md -- task 6 entry under [unreleased]
  3. design/loops/technical.md section 8 -- prompt resolution and token substitution documentation
  4. AGENTS.md -- no changes needed (already documents loop runner and status.py commands)

Findings

1. README.md onboarding section

The new section adds:

  • Quick Start with 3 commands (create, install-schedule, monitor)
  • Tick Cycle diagram (11-step flow summary)
  • Configuration table with all loop.json fields
  • Monitoring commands
  • Halt/Resume commands

Accuracy: All commands and field names match the actual implementation. The configuration table correctly documents use_worktree (not worktree), work_source.kind values (single, audit, backlog), and the role prompt fields.

Completeness: Covers all R7 sub-requirements from the SPEC.

Verdict: PASS

2. CHANGELOG.md

Entry accurately describes all changes: _resolve_prompt, _invoke_harness extension, new prompt files, template updates, new self-improvement template, README section, technical.md section 8, new tests (18), test infrastructure updates.

Verdict: PASS

3. design/loops/technical.md section 8

New "Prompt Resolution and Token Substitution" subsection documents:

  • File search order (loop-local then framework)
  • Content-level token substitution
  • {artifact_content} special handling
  • Temp file write and return path
  • Loop-local override capability

Accuracy: Matches the implementation in _resolve_prompt.

Verdict: PASS

4. Cross-reference check

  • AGENTS.md "Loop runner" bullet references design/loops/technical.md §7 for the tick flow -- still accurate.
  • config.md mentions role-to-prompt binding in loop.json -- still accurate.
  • prompts/ directory now has 3 new files (loop-implement.md, loop-verifier.md, loop-orchestrate.md) -- not listed in any index (there is no prompts/ index file), so no update needed.

Summary

All documentation is accurate, complete, and consistent with the implementation. No doc gaps found.

Verdict: APPROVED