Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
This commit is contained in:
@@ -0,0 +1,53 @@
|
||||
# Implementation: add-goal-mode
|
||||
|
||||
Implements goal-oriented loop extensions per SPEC R1-R8. All changes are in `scripts/loop-runner.py`, `scripts/status.py`, `templates/loops/ci-triage/loop.json`, `design/loops/technical.md`, and `tests/test_goal_mode.py`.
|
||||
|
||||
## Files changed
|
||||
|
||||
- `scripts/loop-runner.py` -- `_find_work` dispatch, `_find_work_audit`, `_find_work_backlog`, `_truncate_tokens`, `_read_task_brief`, `_acceptance_criteria_text`, `_next_hint_text`, new substitution tokens in `cmd_tick`.
|
||||
- `scripts/status.py` -- `--audit --json` mode in `cmd_audit`.
|
||||
- `templates/loops/ci-triage/loop.json` -- explicit `work_source` and `acceptance_criteria` fields.
|
||||
- `design/loops/technical.md` -- schema section updated with `work_source` and `acceptance_criteria`.
|
||||
- `tests/test_goal_mode.py` -- 26 tests covering R1-R8 + regression.
|
||||
|
||||
## R-by-R coverage
|
||||
|
||||
| Req | Code |
|
||||
|-----|------|
|
||||
| R1 find_work dispatch | `_find_work(state, cfg, loop_path, project_dir)` dispatches on `cfg["work_source"]["kind"]`; missing/unknown falls back to `"single"` with WARNING log |
|
||||
| R2 audit work_source | `_find_work_audit` calls `status.py --audit --json`, sorts by severity (high>med>low), uses violation `task` or creates one via `--create-task` |
|
||||
| R3 backlog work_source | `_find_work_backlog` reads `design/<area>/BACKLOG.md`, picks topmost `- [ ]` line, slugifies the `**bold**` heading |
|
||||
| R4 verifier tokens | `{task_brief}`, `{acceptance_criteria}`, `{next_hint}` added to extras dict in `cmd_tick` implement/verify invocations; substituted via `_substitute` |
|
||||
| R5 truncate_tokens | `_truncate_tokens(text, max_tokens)` -- 4 chars/token heuristic, appends ` ...[truncated]` marker; task_brief=4000, acceptance=2000, next_hint=1000 |
|
||||
| R6 next_hint loop | `_next_hint_text(state)` reads `state["last_verdict"]["next_hint"]`; empty on first tick / after approve; fed into both implement and verify |
|
||||
| R7 loop.json schema | ci-triage template updated; technical.md schema section updated |
|
||||
| R8 audit --json | `cmd_audit` in status.py: when `--json`, emits `{"violations":[...], "loops":[...], "total_tasks":N, "untracked_tasks":N}` as single JSON line |
|
||||
|
||||
## Key design decisions
|
||||
|
||||
- `_find_work` returns `(task, skip_reason)` tuple; `skip_reason` is `None` when work found, `"no_current_task"` for single-with-null, `"no_work"` for audit/backlog with no items.
|
||||
- `_find_work_audit` creates tasks via `status.py --create-task` when a violation has no associated task; slug derived from `_slugify(message)`.
|
||||
- `_find_work_backlog` maps `**bold-name**` in checkbox line directly to task name; falls back to slugifying the line body if no bold marker.
|
||||
- Token substitution only applies when the harness command template contains the placeholder; prompts that omit `{task_brief}` etc. are unaffected.
|
||||
- `--audit --json` output is a single JSON line on stdout, parseable by `_run_json` (which takes the last line).
|
||||
|
||||
## Tests (`tests/test_goal_mode.py`)
|
||||
|
||||
26 tests across 8 classes; all `subprocess.run` calls stubbed via monkeypatch.
|
||||
|
||||
- `TestFindWorkDispatch` (3): single work_source; missing work_source falls back to single; unknown kind warns and falls back.
|
||||
- `TestAuditWorkSource` (4): picks highest severity; creates task when no task field; skips when no violations; uses work_source.project override.
|
||||
- `TestBacklogWorkSource` (3): picks top unchecked item; skips when empty; uses area path.
|
||||
- `TestVerifierTokens` (4): task_brief from RESEARCH.md; acceptance_criteria from loop.json list; next_hint from last_verdict; missing tokens leave prompt intact.
|
||||
- `TestTruncateTokens` (3): short text unchanged; long text capped with marker; empty returns empty.
|
||||
- `TestNextHintFeedback` (2): hint fed into next tick; first tick has empty hint.
|
||||
- `TestLoopJsonSchemaAdditions` (3): ci-triage template has work_source; has acceptance_criteria; create_loop preserves acceptance_criteria.
|
||||
- `TestAuditJson` (3): emits violations array; includes loops block; pickable by runner _run_json.
|
||||
- `TestRegressionBackwardCompat` (1): existing single loop with no work_source/acceptance_criteria ticks unchanged.
|
||||
|
||||
## Verification
|
||||
|
||||
- `python3 -m py_compile scripts/loop-runner.py scripts/status.py` -- PASS
|
||||
- `python3 -m pytest tests/test_goal_mode.py -v` -- 26 passed
|
||||
- `python3 -m pytest tests/ -q` -- 354 passed (328 baseline + 26 new)
|
||||
- `bash -n scripts/*.sh` -- no shell changes
|
||||
Reference in New Issue
Block a user