Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

4.9 KiB
Raw Permalink Blame History

SPEC — runnable-test-suite

Goal

Make python3 -m pytest tests/ -v pass from a clean checkout of ~/.automaton, with deterministic Python deps pinned in the repo and documentation that reflects the actual interpreter that ships on the user's machine.

The prime symptom that proves nothing at all runs today: a fresh clone executes python (per AGENTS.md) and silently fails because macOS only ships python3, and even with the right interpreter the suite fails on No module named pytest.

Requirements (numbered)

  1. Add requirements.txt at the repo root pinning pytest (lowest version that supports the syntax used in tests/, which is plain fixtures and tmp_path — pytest ≥ 7.0). No other third-party deps may be added.
  2. Provide a venv-based install path: a one-line install in scripts/install.sh (or a new snippet) that creates .venv/ and pip install -r requirements.txt. Must not require sudo and must not pollute the system Python.
  3. Make python3 -m pytest tests/ -v exit 0 from a clean checkout after pip install -r requirements.txt (no venv required — system pip3 install -r requirements.txt must also work).
  4. Fix every Python file under the repo that fails python3 -m py_compile (currently clean, but must stay clean).
  5. Replace every bare python invocation in documentation and prompts with python3 so the documented commands actually run on a stock macOS without a shim.
    • AGENTS.md lines 58, 61, 64, 70
    • README.md lines 203, 206, 209, 212, 213, 216, 219, 222, 225, 228, 229, 230, 242, 245, 324, 338, 353
    • automaton/dashboard/README.md lines 11, 18, 21, 24
    • prompts/orchestrate.md lines 28, 29, 30, 36
    • Any other python (bare) reference found by rg AFTER the first pass
  6. Do NOT change python references inside shell scripts that already invoke #!/usr/bin/env python3 shebangs or that explicitly resolve via command -v. Only fix bare python commands that shell out (none expected in scripts/ after audit, but verify).
  7. Add a CI step note to CHANGELOG.md under [unreleased] documenting the new requirements.txt and the python3 requirement.
  8. Update AGENTS.md "Build & Test Commands" section to reference requirements.txt and use python3 consistently.

Acceptance criteria

Each must pass from a fresh clone with only stock macOS CommandLineTools + pip3:

  1. pip3 install -r requirements.txt succeeds.
  2. python3 -m pytest tests/ -v exits 0 with N passed (N ≥ 1) and zero error lines.
  3. python3 -m py_compile automaton/**/*.py automaton/dashboard/**/*.py scripts/*.py exits 0.
  4. bash -n scripts/*.sh exits 0.
  5. rg -n "^python |\"python " AGENTS.md README.md automaton/dashboard/README.md prompts/orchestrate.md returns zero matches for a bare python command.
  6. Following the install instructions in AGENTS.md verbatim, a new contributor can run the test suite within 60 seconds of clone.

Success contract (streak)

Per the goal mode this task derives from, "done" requires 10 consecutive clean python3 -m pytest tests/ -v runs in a row without any edit between runs. A single failure resets the counter. The cap on attempts is 5; on hitting the cap, stop and report.

Constraints / non-goals

  • No new dependencies beyond pytest. Do not add pytest-cov, pytest-mock, tox, etc.
  • No virtualenv vendoring. The user creates .venv themselves if they want isolation; system pip3 install -r requirements.txt must also work.
  • No changes to existing test logic. If a test is genuinely broken (not just import-failing because pytest is missing), STOP and report — do not patch the test to make it pass. That is the anti-spin rule from #9 in the source article.
  • Do not touch any file under tasks/ (per-framework tasks are state, not source).
  • Do not modify status.py, vram_detect.py, or any other runtime script's behavior. Only documentation and config files change.
  • No Docker, no conda, no pyenv requirements. Stock python3 + pip3 only.
  • VRAM-aware scoping: this task fits in ONE sub-task (~9k peak context budget on this 32GB-RAM / no-GPU machine with Model: auto). Do not decompose further. Sub-tasks would exceed the budget on overhead alone.
  1. Create requirements.txt with pytest==7.4.4 (last 7.x; works on Python 3.9+).
  2. pip3 install -r requirements.txt locally and run the suite; capture every failure.
  3. For each failure, decide: import/install issue (fix dep) vs. real code bug (report, do not patch test).
  4. Sweep python → python3 in docs/prompts with edit batching.
  5. Add install snippet to scripts/install.sh (idempotent; only if .venv doesn't exist).
  6. Add a one-line test smoke-check at the end of install.sh: python3 -m pytest tests/ -q || echo "tests deferred".
  7. Update CHANGELOG.md [unreleased].
  8. Run the streak verifier: 10× python3 -m pytest tests/ -v; stop at first clean streak or 5 attempts.