Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

4.3 KiB

Implementation: add-goal-mode

Implements goal-oriented loop extensions per SPEC R1-R8. All changes are in scripts/loop-runner.py, scripts/status.py, templates/loops/ci-triage/loop.json, design/loops/technical.md, and tests/test_goal_mode.py.

Files changed

  • scripts/loop-runner.py -- _find_work dispatch, _find_work_audit, _find_work_backlog, _truncate_tokens, _read_task_brief, _acceptance_criteria_text, _next_hint_text, new substitution tokens in cmd_tick.
  • scripts/status.py -- --audit --json mode in cmd_audit.
  • templates/loops/ci-triage/loop.json -- explicit work_source and acceptance_criteria fields.
  • design/loops/technical.md -- schema section updated with work_source and acceptance_criteria.
  • tests/test_goal_mode.py -- 26 tests covering R1-R8 + regression.

R-by-R coverage

Req Code
R1 find_work dispatch _find_work(state, cfg, loop_path, project_dir) dispatches on cfg["work_source"]["kind"]; missing/unknown falls back to "single" with WARNING log
R2 audit work_source _find_work_audit calls status.py --audit --json, sorts by severity (high>med>low), uses violation task or creates one via --create-task
R3 backlog work_source _find_work_backlog reads design/<area>/BACKLOG.md, picks topmost - [ ] line, slugifies the **bold** heading
R4 verifier tokens {task_brief}, {acceptance_criteria}, {next_hint} added to extras dict in cmd_tick implement/verify invocations; substituted via _substitute
R5 truncate_tokens _truncate_tokens(text, max_tokens) -- 4 chars/token heuristic, appends ...[truncated] marker; task_brief=4000, acceptance=2000, next_hint=1000
R6 next_hint loop _next_hint_text(state) reads state["last_verdict"]["next_hint"]; empty on first tick / after approve; fed into both implement and verify
R7 loop.json schema ci-triage template updated; technical.md schema section updated
R8 audit --json cmd_audit in status.py: when --json, emits {"violations":[...], "loops":[...], "total_tasks":N, "untracked_tasks":N} as single JSON line

Key design decisions

  • _find_work returns (task, skip_reason) tuple; skip_reason is None when work found, "no_current_task" for single-with-null, "no_work" for audit/backlog with no items.
  • _find_work_audit creates tasks via status.py --create-task when a violation has no associated task; slug derived from _slugify(message).
  • _find_work_backlog maps **bold-name** in checkbox line directly to task name; falls back to slugifying the line body if no bold marker.
  • Token substitution only applies when the harness command template contains the placeholder; prompts that omit {task_brief} etc. are unaffected.
  • --audit --json output is a single JSON line on stdout, parseable by _run_json (which takes the last line).

Tests (tests/test_goal_mode.py)

26 tests across 8 classes; all subprocess.run calls stubbed via monkeypatch.

  • TestFindWorkDispatch (3): single work_source; missing work_source falls back to single; unknown kind warns and falls back.
  • TestAuditWorkSource (4): picks highest severity; creates task when no task field; skips when no violations; uses work_source.project override.
  • TestBacklogWorkSource (3): picks top unchecked item; skips when empty; uses area path.
  • TestVerifierTokens (4): task_brief from RESEARCH.md; acceptance_criteria from loop.json list; next_hint from last_verdict; missing tokens leave prompt intact.
  • TestTruncateTokens (3): short text unchanged; long text capped with marker; empty returns empty.
  • TestNextHintFeedback (2): hint fed into next tick; first tick has empty hint.
  • TestLoopJsonSchemaAdditions (3): ci-triage template has work_source; has acceptance_criteria; create_loop preserves acceptance_criteria.
  • TestAuditJson (3): emits violations array; includes loops block; pickable by runner _run_json.
  • TestRegressionBackwardCompat (1): existing single loop with no work_source/acceptance_criteria ticks unchanged.

Verification

  • python3 -m py_compile scripts/loop-runner.py scripts/status.py -- PASS
  • python3 -m pytest tests/test_goal_mode.py -v -- 26 passed
  • python3 -m pytest tests/ -q -- 354 passed (328 baseline + 26 new)
  • bash -n scripts/*.sh -- no shell changes