- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
4.3 KiB
4.3 KiB
Implementation: add-goal-mode
Implements goal-oriented loop extensions per SPEC R1-R8. All changes are in scripts/loop-runner.py, scripts/status.py, templates/loops/ci-triage/loop.json, design/loops/technical.md, and tests/test_goal_mode.py.
Files changed
scripts/loop-runner.py--_find_workdispatch,_find_work_audit,_find_work_backlog,_truncate_tokens,_read_task_brief,_acceptance_criteria_text,_next_hint_text, new substitution tokens incmd_tick.scripts/status.py----audit --jsonmode incmd_audit.templates/loops/ci-triage/loop.json-- explicitwork_sourceandacceptance_criteriafields.design/loops/technical.md-- schema section updated withwork_sourceandacceptance_criteria.tests/test_goal_mode.py-- 26 tests covering R1-R8 + regression.
R-by-R coverage
| Req | Code |
|---|---|
| R1 find_work dispatch | _find_work(state, cfg, loop_path, project_dir) dispatches on cfg["work_source"]["kind"]; missing/unknown falls back to "single" with WARNING log |
| R2 audit work_source | _find_work_audit calls status.py --audit --json, sorts by severity (high>med>low), uses violation task or creates one via --create-task |
| R3 backlog work_source | _find_work_backlog reads design/<area>/BACKLOG.md, picks topmost - [ ] line, slugifies the **bold** heading |
| R4 verifier tokens | {task_brief}, {acceptance_criteria}, {next_hint} added to extras dict in cmd_tick implement/verify invocations; substituted via _substitute |
| R5 truncate_tokens | _truncate_tokens(text, max_tokens) -- 4 chars/token heuristic, appends ...[truncated] marker; task_brief=4000, acceptance=2000, next_hint=1000 |
| R6 next_hint loop | _next_hint_text(state) reads state["last_verdict"]["next_hint"]; empty on first tick / after approve; fed into both implement and verify |
| R7 loop.json schema | ci-triage template updated; technical.md schema section updated |
| R8 audit --json | cmd_audit in status.py: when --json, emits {"violations":[...], "loops":[...], "total_tasks":N, "untracked_tasks":N} as single JSON line |
Key design decisions
_find_workreturns(task, skip_reason)tuple;skip_reasonisNonewhen work found,"no_current_task"for single-with-null,"no_work"for audit/backlog with no items._find_work_auditcreates tasks viastatus.py --create-taskwhen a violation has no associated task; slug derived from_slugify(message)._find_work_backlogmaps**bold-name**in checkbox line directly to task name; falls back to slugifying the line body if no bold marker.- Token substitution only applies when the harness command template contains the placeholder; prompts that omit
{task_brief}etc. are unaffected. --audit --jsonoutput is a single JSON line on stdout, parseable by_run_json(which takes the last line).
Tests (tests/test_goal_mode.py)
26 tests across 8 classes; all subprocess.run calls stubbed via monkeypatch.
TestFindWorkDispatch(3): single work_source; missing work_source falls back to single; unknown kind warns and falls back.TestAuditWorkSource(4): picks highest severity; creates task when no task field; skips when no violations; uses work_source.project override.TestBacklogWorkSource(3): picks top unchecked item; skips when empty; uses area path.TestVerifierTokens(4): task_brief from RESEARCH.md; acceptance_criteria from loop.json list; next_hint from last_verdict; missing tokens leave prompt intact.TestTruncateTokens(3): short text unchanged; long text capped with marker; empty returns empty.TestNextHintFeedback(2): hint fed into next tick; first tick has empty hint.TestLoopJsonSchemaAdditions(3): ci-triage template has work_source; has acceptance_criteria; create_loop preserves acceptance_criteria.TestAuditJson(3): emits violations array; includes loops block; pickable by runner _run_json.TestRegressionBackwardCompat(1): existing single loop with no work_source/acceptance_criteria ticks unchanged.
Verification
python3 -m py_compile scripts/loop-runner.py scripts/status.py-- PASSpython3 -m pytest tests/test_goal_mode.py -v-- 26 passedpython3 -m pytest tests/ -q-- 354 passed (328 baseline + 26 new)bash -n scripts/*.sh-- no shell changes