- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/ - status.py: add --cleanup-done and --install-cleanup-schedule commands - Add scripts/automaton-cleanup.sh for periodic task archiving - Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders - .rules.md: add Self-Documenting UI Names rule - New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
3.4 KiB
ADVERSARIAL_BUG_REPORT: add-loop-templates-onboarding
Methodology
Targeted attack on the weakest points of the implementation:
- Path traversal via
prompt_ref - Token injection via
extrasvalues - Race condition on
outputs/directory - Large file DoS via
{artifact_content} - Unicode/encoding edge cases
- Concurrent ticks writing to the same
outputs/dir
Findings
Attack 1: Path traversal via prompt_ref -- NOT VULNERABLE
_resolve_prompt constructs candidate paths as loop_path / prompt_ref and AUTOMATON_DIR / "prompts" / prompt_ref. If prompt_ref were "../../etc/passwd", Path / "../../etc/passwd" would resolve to a path outside the loop dir. However, prompt_ref comes from loop.json roles.*.prompt, which is a trusted config file written by the user/framework. An attacker who can write loop.json already has full code execution via harness.command. No additional risk.
Verdict: NOT VULNERABLE (trusted input)
Attack 2: Token injection via extras values -- NOT VULNERABLE
If task_brief contained {task_brief}, the str(value) substitution would not cause infinite recursion because content.replace is a single-pass operation. The substituted value is inserted as-is, and no further substitution is applied to the result. No injection vector.
Verdict: NOT VULNERABLE
Attack 3: Race condition on outputs/ directory -- NOT EXPLOITABLE
out_dir.mkdir(parents=True, exist_ok=True) is atomic. If two ticks run concurrently (which the scheduler should prevent, but could happen in daemon mode with a bug), they would write to different files (tickN-<role>-prompt.md where N differs). The only shared state is the directory itself, and mkdir(exist_ok=True) handles that. The .state.loop write is atomic (tmp+rename), so iteration_count won't be corrupted.
Verdict: NOT EXPLOITABLE (different tick numbers produce different file paths)
Attack 4: Large file DoS via {artifact_content} -- ACCEPTED RISK
If the implement artifact is very large (e.g. 10MB), {artifact_content} reads the entire file into memory and substitutes it into the prompt. This could produce a prompt that exceeds the model's context window. However, the runner already has a _truncate_tokens function (from task 4) that caps task_brief at 4k tokens, acceptance_criteria at 2k, and next_hint at 1k. The {artifact_content} token is NOT truncated, which is by design -- the verifier needs to see the full artifact to grade it. The 16k context floor gate (D13) catches undersized contexts before the harness is invoked. For oversized contexts, the harness's own context management handles it.
Verdict: ACCEPTED RISK (mitigated by context floor gate and harness-side context management)
Attack 5: Unicode/encoding edge cases -- NOT VULNERABLE
Path.read_text() and Path.write_text() use UTF-8 by default on all platforms. The str(value) conversion handles all Python string types. No encoding issues found.
Verdict: NOT VULNERABLE
Attack 6: Concurrent ticks writing to same outputs/ dir -- NOT EXPLOITABLE
Same as Attack 3. Different tick numbers produce different file paths. The .state.loop atomic write prevents iteration_count corruption.
Verdict: NOT EXPLOITABLE
Summary
No exploitable vulnerabilities found. All attack surfaces are either mitigated by existing controls (context floor gate, atomic state writes, trusted input assumption) or produce no harmful behavior.
Verdict: CLEAN