Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
This commit is contained in:
@@ -0,0 +1 @@
|
||||
complete
|
||||
@@ -0,0 +1,22 @@
|
||||
# Implementation: Harden Dashboard Security
|
||||
|
||||
## Summary
|
||||
- Added CORS headers (`Access-Control-Allow-Origin`, `Methods`, `Headers`) to all API responses via `_send_json()` and `_send_error()`
|
||||
- Added `do_OPTIONS` handler for CORS preflight requests
|
||||
- Added `X-Content-Type-Options: nosniff` header to all responses
|
||||
- Added `MAX_POST_BODY = 65536` (64KB) content-length limit on POST review endpoint
|
||||
- Added `MAX_REVIEW_COMMENT_LENGTH = 4096` character limit on review comments
|
||||
- Replaced inline `onclick` handlers in review buttons with `data-task`/`data-status` attributes + event delegation
|
||||
- Applied `escapeHtml()` to `task.display_name` in `renderTaskCard()`
|
||||
- Filesystem task name validation was already implemented in `fix-verdict-parsing` (R2 of this SPEC is done)
|
||||
|
||||
## Changes
|
||||
- `automaton/dashboard/ui/app.py`: Added CORS headers, `do_OPTIONS`, content-length bounds, comment truncation
|
||||
- `automaton/dashboard/html/dashboard.js`: Replaced onclick handlers with data attributes, escaped display_name
|
||||
|
||||
## Test Results
|
||||
119 passed in 0.08s (full suite)
|
||||
Dashboard starts and serves correct CORS headers on all API responses
|
||||
|
||||
## Blockers
|
||||
None
|
||||
@@ -0,0 +1,57 @@
|
||||
# Harden Dashboard Security
|
||||
|
||||
## Goal
|
||||
|
||||
Close the security gaps identified by the adversarial audit: missing CORS headers, filesystem-sourced task names that bypass validation, and unbounded content-length handling on POST.
|
||||
|
||||
## Requirements
|
||||
|
||||
### R1. Add CORS headers
|
||||
|
||||
The dashboard serves no CORS headers. When bound to `0.0.0.0` (documented in `__main__.py`), any webpage can call the API — including approving/rejecting tasks via POST.
|
||||
|
||||
**Fix**: In `DashboardHandler._send_json()` and `_send_error()`, add:
|
||||
- `Access-Control-Allow-Origin: *` (or configurable via `--cors-origin`)
|
||||
- `Access-Control-Allow-Methods: GET, POST, OPTIONS`
|
||||
- `Access-Control-Allow-Headers: Content-Type`
|
||||
- Handle `OPTIONS` preflight requests for the review endpoint
|
||||
|
||||
### R2. Validate filesystem-sourced task names
|
||||
|
||||
`discover_tasks()` at `task.py:248` reads directory names directly from `iterdir()`. The `_validate_task_name` regex only applies to API path parsing. A task directory created via `mkdir` with special characters (e.g., quotes, HTML) will be served to the JS client, which injects names into `onclick` attributes and `innerHTML`.
|
||||
|
||||
**Fix**: In `discover_tasks()`, skip directories whose names contain characters outside `[A-Za-z0-9_-]`. Log a warning for invalid names.
|
||||
|
||||
### R3. Add content-length bound check on POST regardless of R2 from wire-dashboard-config
|
||||
|
||||
Even if the caching task isn't done yet, add a quick defensive check:
|
||||
- If `Content-Length` header > `MAX_POST_BODY`, return 413
|
||||
- If `Content-Length` header is missing or <= 0, return 400
|
||||
|
||||
### R4. Escape task names in JS HTML injection points
|
||||
|
||||
In `dashboard.js:renderDetail()`, `renderTaskCard()`, and `renderTimeline()`, task names are interpolated into HTML. While R2 prevents most dangerous names, defense in depth requires:
|
||||
|
||||
- Use `escapeHtml()` on `task.display_name` before injection
|
||||
- Use `data-*` attributes instead of `onclick` for review buttons (pass task name via `dataset`)
|
||||
|
||||
### R5. Add `X-Content-Type-Options: nosniff` header
|
||||
|
||||
All responses should include `X-Content-Type-Options: nosniff` to prevent MIME type sniffing.
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] All API responses include `Access-Control-Allow-Origin` header
|
||||
- [ ] `OPTIONS /api/task/{name}/review` returns 200 with appropriate CORS headers
|
||||
- [ ] Task directory named `task-with'quote` is excluded from `discover_tasks()` output
|
||||
- [ ] Task directory named `valid-task-123` is included
|
||||
- [ ] POST with `Content-Length: 1000000` returns 413 regardless of caching task status
|
||||
- [ ] `escapeHtml()` applied to `display_name` in all JS interpolation points
|
||||
- [ ] Review buttons use `data-task` attribute instead of inline `onclick`
|
||||
- [ ] All responses include `X-Content-Type-Options: nosniff`
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- Not adding authentication (the dashboard is a local single-user tool)
|
||||
- Not adding HTTPS (out of scope for a dev tool)
|
||||
- Not rate-limiting (single-user, single-threaded server)
|
||||
@@ -0,0 +1,21 @@
|
||||
# Verdict: harden-dashboard-security
|
||||
|
||||
## Status: PASS
|
||||
**Completion Date**: 2026-06-14
|
||||
|
||||
## Summary
|
||||
Closed dashboard security gaps: CORS headers on all API responses, POST content-length bounds (64KB), review comment length limits (4096 chars), XSS defense via escapeHtml on display names and data attributes instead of inline onclick, filesystem task name validation inherited from fix-verdict-parsing.
|
||||
|
||||
## Findings
|
||||
- All 119 tests pass (including 7 new security/CORS tests)
|
||||
- CORS headers present on all JSON responses and OPTIONS preflight
|
||||
- X-Content-Type-Options: nosniff on all responses
|
||||
- Review POST rejects Content-Length > 65536 with 413
|
||||
- Review comment truncated to 4096 characters
|
||||
- Task names with special characters are excluded from discover_tasks output (implemented in fix-verdict-parsing)
|
||||
|
||||
## Tasks for Review / Tie-Breaks
|
||||
- None
|
||||
|
||||
## Score
|
||||
+10
|
||||
Reference in New Issue
Block a user