Files
automaton/tasks/complete/fix-context-sizing/ADVERSARIAL_BUG_REPORT.md
T
Lap Tran 4a2301b077
CI / build (push) Has been cancelled
Archive completed tasks, add cleanup commands, self-documenting dashboard UI
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2026-06-24 22:43:33 -04:00

35 lines
2.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Adversarial Bug Report: fix-context-sizing
## Status: NO_NEW_DEFECTS
Adversarial review applied the 5 attack vectors from `design/loops/functional.md` §6 to confirm the changes don't introduce enforcement gaps:
A1. Concurrent state divergence
A2. Single-session harness loops can merge 2 sessions
A3. VM/CI environment with no GPU detection
A4. User types "loop mode" instead of "--loop-mode"
A5. Model auto-picks "auto" through config.md but isn't detected
Note: the loop system itself ( brakes, runner, verifier) is tasks 2–4. This task only ships the *foundation* the loop runner consumes. Adversarial focus is therefore: (a) has the foundation been honestly graded, (b) can it break the existing enforcement layer, (c) does it lie in a way that loops would silently accept bad budgets.
## Attacks
### A1: Concurrent state divergence
`vram_detect.py` is read-only w.r.t. task state. No `.state` file mutation. No enforcement-layer coupling. Safe by design.
### A2: Single-session harness fails to detect two sessions
Not in scope for this task. Roles and harness invocation are task 6's `templates/loops/`. This task only adds the `## Loop Role Models` documentation block to `config.md`; no logic affects session binding.
### A3: VM/CI environment, GPU detection returns 0
The fallback chain in `recommend_context` already handles `gpu_vram_gb == 0` (falls through to RAM or model context). The new `--loop-mode` floor check correctly refuses when `max_peak_kb < 16_000`. Manually exercised logic with `gpu_vram_gb=0, ram_gb=8, model_context_kb=0`; `max_peak_kb` computation flows through RAM branch (8 * 750 = 6000), minus overhead, * 75% = under 16k → loop-mode refuses. Behavior correct.
### A4: User misspells conf
Mangled flag is rejected by argparse; not silent. Confirmed `--help` shows the flag; argparse errors on unknown flag. Safe.
### A5: Model "auto" in config.md
Existing flow already handles "auto" by returning None from `_parse_config_model` value check (line 454, 458). When user has left `Model: auto` AND no `Override context window`: `detect_model_context` proceeds to attempt API config probing and ollama probe. If both fail, returns 0, and `--loop-mode` refuses with the exact D13 message. Non-loop mode warns. Matches SPEC intent.
## New Defects
None.
## Adversarial Verdict
Foundation holds. All 5 attacks correctly produce refuse/warn behavior or are out-of-scope for this task. Ready to ship.