- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
49 lines
2.4 KiB
Markdown
49 lines
2.4 KiB
Markdown
# ADVERSARIAL_BUG_REPORT: add-self-improvement-loop
|
|
|
|
## Methodology
|
|
|
|
Targeted attack on:
|
|
1. Shell injection via `$FRAMEWORK_DIR`
|
|
2. Race condition between install.sh and update.sh
|
|
3. Loop creation failure cascading to install failure
|
|
4. Schedule installation on unsupported platforms
|
|
5. Template path traversal
|
|
|
|
## Findings
|
|
|
|
### Attack 1: Shell injection via `$FRAMEWORK_DIR` -- NOT VULNERABLE
|
|
|
|
`$FRAMEWORK_DIR` is set to `$HOME/.automaton` at the top of both scripts. It is not derived from user input. The `--project "$FRAMEWORK_DIR"` argument is passed as a single quoted argument to `python3`, so no shell expansion occurs inside the Python process. No injection vector.
|
|
|
|
**Verdict:** NOT VULNERABLE
|
|
|
|
### Attack 2: Race condition between install.sh and update.sh -- NOT EXPLOITABLE
|
|
|
|
If a user runs `install.sh` and `update.sh` concurrently (which would be unusual), both might try to create the loop simultaneously. `--create-loop` checks `if loop_path.exists()` and returns rc=2 if it exists. The `mkdir(parents=True)` in `cmd_create_loop` is not atomic, but the `.state.loop` write is atomic (tmp+rename). Worst case: one script gets rc=2 and `|| true` swallows it. No data corruption.
|
|
|
|
**Verdict:** NOT EXPLOITABLE
|
|
|
|
### Attack 3: Loop creation failure cascading -- NOT VULNERABLE
|
|
|
|
Both `--create-loop` and `--install-schedule` are followed by `|| true`. If either fails, the script continues. The `.venv` setup and pip install at the end of `install.sh` are outside the `else` block and run regardless. The framework works without the loop.
|
|
|
|
**Verdict:** NOT VULNERABLE
|
|
|
|
### Attack 4: Schedule installation on unsupported platforms -- HANDLED
|
|
|
|
`--install-schedule` handles platform dispatch internally (Darwin -> launchd, Linux -> cron, Windows -> schtasks). On an unknown platform, it prints an error and returns non-zero, which `|| true` swallows. The loop is created but not scheduled; the user can manually run `--mode tick` or `--mode daemon`.
|
|
|
|
**Verdict:** HANDLED
|
|
|
|
### Attack 5: Template path traversal -- NOT VULNERABLE
|
|
|
|
`--from-template self-improvement` is a fixed string in both scripts. `cmd_create_loop` constructs the template path as `AUTOMATON_DIR / "templates" / "loops" / template`. The template name is not user-supplied in this context.
|
|
|
|
**Verdict:** NOT VULNERABLE
|
|
|
|
## Summary
|
|
|
|
No exploitable vulnerabilities found. All attack surfaces are mitigated by trusted input, `|| true` non-fatal behavior, and atomic state writes.
|
|
|
|
**Verdict: CLEAN**
|