Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

2.4 KiB

ADVERSARIAL_BUG_REPORT: add-self-improvement-loop

Methodology

Targeted attack on:

  1. Shell injection via $FRAMEWORK_DIR
  2. Race condition between install.sh and update.sh
  3. Loop creation failure cascading to install failure
  4. Schedule installation on unsupported platforms
  5. Template path traversal

Findings

Attack 1: Shell injection via $FRAMEWORK_DIR -- NOT VULNERABLE

$FRAMEWORK_DIR is set to $HOME/.automaton at the top of both scripts. It is not derived from user input. The --project "$FRAMEWORK_DIR" argument is passed as a single quoted argument to python3, so no shell expansion occurs inside the Python process. No injection vector.

Verdict: NOT VULNERABLE

Attack 2: Race condition between install.sh and update.sh -- NOT EXPLOITABLE

If a user runs install.sh and update.sh concurrently (which would be unusual), both might try to create the loop simultaneously. --create-loop checks if loop_path.exists() and returns rc=2 if it exists. The mkdir(parents=True) in cmd_create_loop is not atomic, but the .state.loop write is atomic (tmp+rename). Worst case: one script gets rc=2 and || true swallows it. No data corruption.

Verdict: NOT EXPLOITABLE

Attack 3: Loop creation failure cascading -- NOT VULNERABLE

Both --create-loop and --install-schedule are followed by || true. If either fails, the script continues. The .venv setup and pip install at the end of install.sh are outside the else block and run regardless. The framework works without the loop.

Verdict: NOT VULNERABLE

Attack 4: Schedule installation on unsupported platforms -- HANDLED

--install-schedule handles platform dispatch internally (Darwin -> launchd, Linux -> cron, Windows -> schtasks). On an unknown platform, it prints an error and returns non-zero, which || true swallows. The loop is created but not scheduled; the user can manually run --mode tick or --mode daemon.

Verdict: HANDLED

Attack 5: Template path traversal -- NOT VULNERABLE

--from-template self-improvement is a fixed string in both scripts. cmd_create_loop constructs the template path as AUTOMATON_DIR / "templates" / "loops" / template. The template name is not user-supplied in this context.

Verdict: NOT VULNERABLE

Summary

No exploitable vulnerabilities found. All attack surfaces are mitigated by trusted input, || true non-fatal behavior, and atomic state writes.

Verdict: CLEAN