Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

5.5 KiB
Raw Permalink Blame History

Implementation: add-status-brakes

Implements SPEC.md R1–R10. All new code lives in scripts/status.py (loop extensions) plus a new test file tests/test_status_brakes.py and a minimal loop template at templates/loops/ci-triage/loop.json.

Surface added (R1–R10)

Req CLI surface Behavior
R1 n/a .state.loop schema v1 with 13 default fields; written atomically via tmp+rename
R2 --create-loop NAME [--from-template T] Refuses non-kebab, duplicates, unknown template; patches name into copied loop.json; seeds empty .state.log
R3 --version; --approve --loop NAME --version reads ## Framework Version from config.md; --approve --loop is the only way to clear a halt (D4); increments resumed_count
R4 --can-continue NAME [--json] Cheap status probe: ok := status == "running"
R5 --check-gate NAME [--json] Runs 6 gates in order; first failure halts the loop and emits structured verdict
R6 --install-schedule NAME [--interval S] Generates run-tick.sh/.bat; installs launchd plist / crontab block / schtasks unit per platform.system(); --pause-loop best-effort disables the unit
R7 --can-edit --loop NAME [--loop-worktree] --file P Checks file against loop's blast_radius.file_scope; refuses files outside project/framework root
R8 --transition extension Refuses if a HALTED loop owns the task (_loop_owning_task scan); points user at --approve --loop
R9 --audit Cat-6 block; --loop-list Reuses _audit_loops_block; runs even when no tasks exist
R10 .state.log tick trail Every state-changing op appends an ISO-timestamped line; tests assert PAUSED/RESUMED/APPROVED/HALT are all logged

Gate order (R5)

gate_loop_status      -> not running -> halt w/ existing halt_reason
gate_iterations       -> iteration_count >= max_iterations -> iterations_exhausted
gate_budget           -> spent_usd >= max_budget_usd        -> budget_exhausted (remote-only, informational)
gate_task_phase       -> current_task in human_intervention -> human_intervention
gate_worktree_drift   -> changed files outside file_scope    -> drift_detected
gate_score_plateau    -> score_history flat across window    -> verifier_failed

First failure wins. Halt is written atomically; schedule is best-effort disabled.

Helper functions added (scripts/status.py, before def main())

  • LOOP_* constants (states, halts, schema version, file names)
  • _loops_dir, _loop_dir, _all_loop_dirs
  • _read_state_loop, _write_state_loop, _initial_state_loop, _read_loop_config
  • _append_tick_log, _loop_untracked_hint
  • _halt_loop, _disable_schedule, _enable_schedule
  • _loop_owning_task (R8 ownership scan)
  • _gate_* (6 gate functions)
  • _loop_max_iterations
  • _task_phase_for_loop
  • cmd_create_loop, cmd_install_schedule, cmd_pause_loop, cmd_resume_loop
  • cmd_approve_loop (R3 halt-clear)
  • cmd_check_gate, cmd_can_continue
  • cmd_loop_list, cmd_version
  • cmd_can_edit_loop (R7 worktree scope)

Existing functions extended

  • cmd_can_edit — early hook: if args.loop, delegate to cmd_can_edit_loop.
  • cmd_transition — R8 halt-refusal inserted after _require_state; _loop_owning_task scan.
  • cmd_audit — _audit_loops_block(args) helper called twice (early-return empty-tasks path + main path); Cat-6 header always printed.

Argparse additions (main())

--create-loop, --from-template, --install-schedule, --interval, --pause-loop, --resume-loop, --loop, --loop-worktree, --check-gate, --can-continue, --loop-list, --version.

Dispatch order places loop commands before task commands so --approve --loop doesn't fall through to the --task-required cmd_approve.

New file: templates/loops/ci-triage/loop.json

Minimal template used as --create-loop default. Defines brakes.max_iterations=25, score_plateau_window=5, blast_radius.use_worktree=true. Full prompt/template expansion is task 6.

Tests

tests/test_status_brakes.py — 46 tests across 10 classes mirroring R1–R10:

  • TestStateLoopSchema (R1) — default-schema assertions + tick log file presence
  • TestCreateLoop (R2) — kebab/dup/template rejection + name-patching
  • TestVersionAndApprove (R3) — version regex; approve refuses non-halted; clears halted + bumps resumed_count
  • TestCanContinue (R4) — running ok, halted denied, unknown → exit 2
  • TestCheckGate (R5) — fresh-pass, status-halt, iterations-exhausted, iterations-remaining, budget-exhausted, budget-informational, task-phase-halt, score-plateau, short-history-ok, JSON output
  • TestInstallSchedule (R6) — stub generation, default interval from config, unknown-loop rejection
  • TestCanEditLoop (R7) — in-scope allowed, out-of-scope denied, outside-root denied, no-file rejected
  • TestTransitionHaltRefusal (R8) — refused when halted owner, allowed when running owner, allowed when no owner
  • TestAuditAndList (R9) — empty list, populated list, Cat-6 header on empty, halted flag, untracked flag, running-pass
  • TestTickLog (R10) — PAUSED/RESUMED/APPROVED/HALT all logged
  • TestPauseResume — pause sets paused; resume only from paused; halted→approve pointer

Verification

python3 -m py_compile scripts/status.py        # OK
python3 -m pytest tests/test_status_brakes.py -q  # 46 passed
python3 -m pytest tests/ -q                       # 310 passed (was 264 + 46 new)

No existing tests changed. Full suite green.