Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
This commit is contained in:
@@ -0,0 +1 @@
|
||||
complete
|
||||
@@ -0,0 +1,35 @@
|
||||
# Adversarial Bug Report: task-status-reason
|
||||
|
||||
## Summary
|
||||
Deep-dive audit found 3 subtle issues: misleading red border on bug_find cards, redundant verdict parsing, and no status_reason propagation to subtasks.
|
||||
|
||||
## Bugs Found
|
||||
|
||||
### Bug 1: Red error border on bug_find task cards is misleading
|
||||
- **Severity**: Low
|
||||
- **Location**: `automaton/dashboard/html/styles.css` — `.task-card-reason` rule
|
||||
- **Description**: The `.task-card-reason` CSS uses `border-left: 2px solid var(--error)` for all three states (blocked, bug_find, adv_bug_find). But `bug_find` and `adv_bug_find` are normal workflow states, not errors. A green or neutral border would be more appropriate for non-terminal states.
|
||||
- **Reproduction**: Open any task in bug_find or adv_bug_find state — the reason banner shows a red border suggesting something is wrong, when it's expected behavior.
|
||||
- **Suggested Fix**: Use `var(--warning)` for bug_find states and reserve `var(--error)` for blocked only.
|
||||
|
||||
### Bug 2: Redundant verdict parsing on every status_reason access
|
||||
- **Severity**: Low
|
||||
- **Location**: `automaton/dashboard/core/task.py:136-175` — `status_reason` property
|
||||
- **Description**: `status_reason` calls `parse_verdict_status()` again for BLOCKED/DONE tasks, even though `determine_task_state()` already parsed the verdict to classify the state. For DONE, it doesn't re-parse (just returns "Verdict: PASS"), but for BLOCKED it re-reads the verdict content and re-parses. This is redundant but cheap given the small number of tasks.
|
||||
- **Reproduction**: Every time the dashboard renders, BLOCKED tasks trigger a second verdict parse just for the display string.
|
||||
- **Suggested Fix**: Cache the parsed verdict status on the Task object during `discover_tasks()`, e.g., a `_verdict_status` field that `status_reason` can reference instead of re-parsing.
|
||||
|
||||
### Bug 3: Subtask state not visible in detail panel reason
|
||||
- **Severity**: Low
|
||||
- **Location**: `automaton/dashboard/html/dashboard.js:236-241` — subtask list in detail panel
|
||||
- **Description**: Subtasks in the detail panel show only a verdict pass/fail indicator. Unlike parent tasks, blocked subtasks don't display their status_reason in the list. A blocked subtask just shows "✗" next to its name with no explanation of why.
|
||||
- **Suggested Fix**: For blocked subtasks, include the subtask's verdict status (FAIL/NEEDS_REVIEW) in the list item, or hover tooltip with the reason.
|
||||
|
||||
### Bug 4: Fallthrough status_reason never reached, masks missing states
|
||||
- **Severity**: Low
|
||||
- **Location**: `automaton/dashboard/core/task.py:175`
|
||||
- **Description**: The fallthrough `return f"In {self.state.value} phase"` is dead code — every TaskState value is covered by explicit branches. If a new state is added (e.g., INTEGRATION_TEST), no compile-time error occurs and a vague message is shown. Contrast with Python enums which have no exhaustiveness checking.
|
||||
- **Suggested Fix**: Add `# pragma: no cover` or raise/log a warning if a new state goes unhandled.
|
||||
|
||||
## Score
|
||||
+3
|
||||
@@ -0,0 +1,30 @@
|
||||
# Bug Report: task-status-reason
|
||||
|
||||
## Summary
|
||||
Audit of status_reason implementation against SPEC.md found minor issues: a dead-code branch in the status_reason message, and a pre-existing REFEREE state that's never produced.
|
||||
|
||||
## Bugs Found
|
||||
|
||||
### Bug 1: Dead message branch in status_reason for BLOCKED state
|
||||
- **Severity**: Low
|
||||
- **Location**: `automaton/dashboard/core/task.py:149`
|
||||
- **Description**: The message `"Verdict is empty or could not be parsed"` has two scenarios:
|
||||
1. Empty verdict → reachable (content is `""`, correctly returns this message)
|
||||
2. Unparseable verdict → unreachable. When `parse_verdict_status()` returns `None`, the state machine in `determine_task_state()` doesn't classify the task as BLOCKED — it falls through to earlier state checks. So a task is never both BLOCKED and "could not be parsed".
|
||||
- **Suggested Fix**: Change message to `"Verdict is empty"` to accurately reflect the only reachable case.
|
||||
|
||||
### Bug 2: REFEREE state never produced by state machine
|
||||
- **Severity**: Low
|
||||
- **Location**: `automaton/dashboard/core/task.py:174` (fallback `status_reason` line), and `task.py:226-272` (determine_task_state)
|
||||
- **Description**: `TaskState.REFEREE` exists in the enum and in the sort order, but `determine_task_state` never returns it. When a VERDICT.md is present but unparseable, the state machine falls through to earlier artifact checks instead of assigning REFEREE. This means the fallback status reason `"In referee phase"` is unreachable, and tasks with ambiguous verdicts silently show as earlier states (e.g., RESEARCH if only SPEC.md exists alongside an unparseable VERDICT.md).
|
||||
- **Pre-existing**: This predates the task and is not introduced by the implementation, but the status_reason property exposes it because the REFEREE branch is dead code.
|
||||
- **Suggested Fix**: Add `if "VERDICT.md" in artifacts: return TaskState.REFEREE, artifacts` after the terminal-state checks in `determine_task_state()`, before the state machine fallthrough.
|
||||
|
||||
### Bug 3: REFEREE status reason is generic
|
||||
- **Severity**: Low
|
||||
- **Location**: `automaton/dashboard/core/task.py:175`
|
||||
- **Description**: The fallthrough line `return f"In {self.state.value} phase"` would produce generic messages like `"In blocked phase"` or `"In done phase"` for states that are handled above it. This is actually dead code for all defined states since every `TaskState` value is covered by an explicit `if` branch. If a new state is added without adding a status_reason handler, it gets a generic message rather than failing loudly.
|
||||
- **Suggested Fix**: Replace the fallthrough with a clear signal: either raise an error, or explicitly list the catch to alert developers when adding states.
|
||||
|
||||
## Score
|
||||
+5
|
||||
@@ -0,0 +1,24 @@
|
||||
# Documentation Review: task-status-reason
|
||||
|
||||
## Summary
|
||||
No DESIGN.md. Conducting ad-hoc review of code documentation for the status_reason feature.
|
||||
|
||||
## Documentation Completeness
|
||||
- Code documentation: Adequate — `status_reason` property has a docstring. `parse_verdict_status` was already documented. No DESIGN.md or deployment docs produced (proportional to the small change set).
|
||||
- User documentation: Not updated — the dashboard README (`automaton/dashboard/README.md`) doesn't mention the status reason display. The new status_reason field in the API response is also undocumented.
|
||||
- API documentation: No formal API docs exist for the dashboard endpoints.
|
||||
|
||||
## Issues Found
|
||||
|
||||
### Issue 1: Dashboard README not updated for status reason display
|
||||
- **Severity**: Low
|
||||
- **Description**: The dashboard README at `automaton/dashboard/README.md` doesn't document that the detail panel now shows a status reason. This is a user-facing UI change without documentation.
|
||||
- **Suggested Fix**: Add a note to the README under "Detail Panel" section describing the status reason.
|
||||
|
||||
### Issue 2: CHANGELOG entry missing bug fixes
|
||||
- **Severity**: Low
|
||||
- **Description**: The CHANGELOG correctly lists the status_reason feature and revoke buttons, but doesn't mention the REFEREE state fix (a consequence of Bug 2 fix from BUG_REPORT.md).
|
||||
- **Suggested Fix**: Add a Fixed entry for the REFEREE state being unreachable.
|
||||
|
||||
## Score
|
||||
+3
|
||||
@@ -0,0 +1,17 @@
|
||||
# Implementation: Add Status Reason to Task Display
|
||||
|
||||
## Summary
|
||||
Added `status_reason` property to `Task` model providing human-readable explanations for each task state, exposed it in API responses, displayed it prominently in the detail panel and on task cards. Fixed the pending review count to exclude terminal-state tasks. Replaced approve/request-changes buttons with revoke buttons when a review is already submitted.
|
||||
|
||||
## Changes
|
||||
- `automaton/dashboard/core/task.py`: Added `status_reason` property on `Task` with explanations for all 10+ task states
|
||||
- `automaton/dashboard/ui/app.py`: Added `status_reason` to `_serve_tasks()` and `_serve_task()` JSON responses; accepted `"pending"` as valid review status
|
||||
- `automaton/dashboard/html/dashboard.js`: Displayed `status_reason` in detail panel and on task cards; fixed pending count to exclude done/blocked; added conditional button rendering for approved/changes_requested reviews
|
||||
- `automaton/dashboard/html/styles.css`: Added `.detail-status-reason` and `.task-card-reason` styles
|
||||
- `tests/test_task.py`: Added `TestStatusReason` class with 5 tests covering done, blocked, and in-progress states
|
||||
|
||||
## Test Results
|
||||
156 passed
|
||||
|
||||
## Blockers
|
||||
None
|
||||
@@ -0,0 +1,40 @@
|
||||
# Add Status Reason to Task Display
|
||||
|
||||
## Goal
|
||||
Show a human-readable explanation of *why* a task is in its current state, fix the pending review count to exclude terminal-state tasks, and replace approve/request-changes buttons with revoke buttons when a review is already submitted.
|
||||
|
||||
## Requirements
|
||||
|
||||
### R1. Add `status_reason` property to Task
|
||||
Add a `status_reason` property on the `Task` dataclass that derives a human-readable explanation from the task's state, artifacts, and verdict status. For example:
|
||||
- DONE → "Verdict: PASS"
|
||||
- BLOCKED → "Verdict: FAIL — changes required before re-review"
|
||||
- BUG_FIND (with IMPLEMENTATION.md) → "Implementation complete — awaiting bug finding"
|
||||
- BACKLOG → "No artifacts yet — not started"
|
||||
|
||||
### R2. Expose `status_reason` in API responses
|
||||
Include `status_reason` in the JSON response for both `/api/tasks` and `/api/task/{name}`.
|
||||
|
||||
### R3. Display status reason prominently
|
||||
- Show `status_reason` below the status badge in the task detail panel
|
||||
- Show `status_reason` on task cards for blocked/bug_find/adv_bug_find states
|
||||
|
||||
### R4. Fix pending review count
|
||||
The pending review counter in the header currently counts all tasks without a review, including DONE tasks. Fix it to exclude terminal states (done, blocked) so only active tasks count toward pending.
|
||||
|
||||
### R5. Replace review buttons with revoke on submitted reviews
|
||||
When a task already has an approved review, replace the approve button with "Revoke Approval". When changes are requested, replace with "Revoke Changes". Both send status="pending" to reset the review.
|
||||
|
||||
## Acceptance Criteria
|
||||
- [ ] `status_reason` property works for all task states
|
||||
- [ ] `/api/tasks` responses include `status_reason`
|
||||
- [ ] Detail panel shows status reason below the status badge
|
||||
- [ ] Task cards show status reason for blocked/bug_find/adv_bug_find tasks
|
||||
- [ ] Pending review count excludes done/blocked tasks
|
||||
- [ ] Approved tasks show "Revoke Approval" instead of "Approve"
|
||||
- [ ] Backend accepts "pending" as a review status
|
||||
- [ ] All tests pass
|
||||
|
||||
## Non-Goals
|
||||
- Not changing the task state machine — revoking a review only resets the review metadata, not the task state
|
||||
- Not adding buttons to change task state directly (e.g., send back to planning) — that's a larger feature
|
||||
@@ -0,0 +1,19 @@
|
||||
# Verdict: task-status-reason
|
||||
|
||||
## Status: PASS
|
||||
**Completion Date**: 2026-06-14
|
||||
|
||||
## Summary
|
||||
Added status_reason to Task model with human-readable explanations for each state, fixed pending review count to exclude terminal-state tasks, and replaced approve/request-changes buttons with revoke buttons on submitted reviews.
|
||||
|
||||
## Findings
|
||||
- All 156 tests pass (5 new status_reason tests)
|
||||
- status_reason covers all 10+ task states with specific messages derived from artifacts
|
||||
- Backend accepts "pending" as review status for revocations
|
||||
- Pending count now only shows non-terminal tasks needing review
|
||||
|
||||
## Tasks for Review / Tie-Breaks
|
||||
- None
|
||||
|
||||
## Score
|
||||
+10
|
||||
Reference in New Issue
Block a user