Files
automaton/tasks/complete/add-goal-mode/BUG_REPORT.md
T
Lap Tran 4a2301b077
CI / build (push) Has been cancelled
Archive completed tasks, add cleanup commands, self-documenting dashboard UI
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2026-06-24 22:43:33 -04:00

2.8 KiB

BUG_REPORT: add-goal-mode

Probed goal-mode work sources, token substitution, and audit --json against edge cases.

Bugs found

None blocking. Informational observations below.

Observations (non-blocking)

O1 -- _find_work_audit create-task subprocess is fire-and-forget

When a violation has no task field, the runner calls status.py --create-task <slug> with timeout=15 and swallows all exceptions. If the create-task fails (e.g. disk full, permission error), the runner returns the slug anyway. The next tick's --audit will see the same violation (still no task dir) and try again. Self-healing on next tick. Accepted for v1.

O2 -- _find_work_backlog bold-marker regex is strict

The regex \*\*([A-Za-z0-9._-]+)\*\* requires the bold text to be a valid slug (alphanumerics, dots, hyphens, underscores only). A backlog item like - [ ] **fix user auth** would fail the regex and fall back to _slugify("fix user auth") -> fix-user-auth. This is correct behavior but worth noting: the bold marker is a convention, not a requirement. Accepted.

O3 -- --audit --json violations lack resolved: true entries

The audit collector only emits unresolved violations (those with actual defects). Resolved violations are not included in the JSON output. This is correct for the runner's use case (it filters on not v.get("resolved", False) anyway), but a consumer expecting a full audit history would need the human-readable --audit output instead. Accepted.

O4 -- Token substitution tests require custom harness command

The R4 tests (test_task_brief_substituted_from_research, test_acceptance_criteria_substituted_from_loop_json_list, test_next_hint_substituted_from_last_verdict) use a custom harness.command that includes the token placeholders. The default harness command (opencode run --prompt-file {prompt} --cwd {cwd}) does not contain {task_brief} etc., so the tokens are only useful when a loop config explicitly adds them to its harness command. This is by design (SPEC R4: "tokens absent from the prompt stay literal"). The actual prompt files (task 6) will need to either reference these tokens or the harness command will need to pass them as CLI args. Accepted.

O5 -- _truncate_tokens marker length can exceed budget by 1

The marker is ...[truncated] (14 chars with leading space). The code does text[:char_budget - len(_TRUNCATE_MARKER)] + marker. If char_budget is smaller than len(_TRUNCTATE_MARKER), the slice goes negative and Python returns the whole string (not empty). For max_tokens=1 (budget=4), the result would be the full text + marker. This only happens with absurdly small token budgets (the real caps are 1000-4000). Not blocking. Noted for v1.1 hardening: clamp char_budget to len(marker) + 1 minimum.

Verdict

PASS -- no blocker bugs. All observations are accepted trade-offs or v1.1 hardening items.