Files
automaton/tasks/complete/fix-context-sizing/SPEC.md
T
Lap Tran 4a2301b077
CI / build (push) Has been cancelled
Archive completed tasks, add cleanup commands, self-documenting dashboard UI
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2026-06-24 22:43:33 -04:00

164 lines
11 KiB
Markdown

# Fix Context Sizing
Tier 1 context-sizing fixes — the foundational layer that `add-status-brakes` and downstream loop tasks consume. Per `design/loops/technical.md` §11 and the locked Tier 1 list.
## Goal
Make the framework's context-budget reporting honest, single-headroom-applied, and machine-readable with a hard floor. Today `vram_detect.py` lies: it reports fake 8k/6k defaults when the actual budget is zero, silently applies headroom twice, and never refuses to run on an unknown model. The loop runner (task `add-loop-runner`) needs accurate, authoritative numbers — its 16k floor check (D13) is meaningless against fabricated defaults.
## Requirements
### R1. Remove double-headroom application in `recommend_context`
`scripts/vram_detect.py:618-656` applies headroom twice:
- Line 642 / 644 / 648: applies `(100 - headroom_pct) // 100` while constructing `recommended_kb` from `vram_context_kb` / `model_context_kb` / `ram_context_kb`.
- Line 654: applies `(100 - headroom_pct) // 100` again when deriving `max_peak_kb = net_kb * (100 - headroom_pct) // 100`.
Net effect: `max_peak_kb` is discounted by `headroom_pct` *twice*, so a 25% headroom becomes a 44% reduction.
**Fix**: Restructure `recommend_context` so headroom is applied **exactly once**. Build `recommended_kb` as the raw budget (no `* (100 - headroom_pct) // 100` at lines 642, 644, 648), subtract overhead, then apply headroom once to derive `max_peak_kb`:
```python
def recommend_context(...) -> tuple[int, int, int]:
headroom_pct = int(config.get("headroom_pct", DEFAULT_HEADROOM_PCT))
if not config.get("auto_detect", True):
target_kb = int(config.get("target_context_kb", 0))
max_peak_kb = int(config.get("max_peak_kb", 0))
if target_kb > 0:
if max_peak_kb == 0 and headroom_pct > 0:
max_peak_kb = target_kb * (100 - headroom_pct) // 100
return headroom_pct, target_kb, max_peak_kb
recommended_kb = 0
if gpu_vram_gb >= 4:
recommended_kb = gpu_vram_gb * 2000
elif model_context_kb > 0:
recommended_kb = model_context_kb
else:
recommended_kb = ram_gb * 750
net_kb = recommended_kb - overhead_tokens
max_peak_kb = net_kb * (100 - headroom_pct) // 100
return headroom_pct, net_kb, max_peak_kb
```
Manual-override branch unchanged (it already applies headroom once via `max_peak_kb = target_kb * (100 - headroom_pct) // 100`).
### R2. Stop lying about zero/negative budgets
`scripts/vram_detect.py:651`:
```python
net_kb = max(0, recommended_kb - overhead_tokens)
```
and `:698-699`:
```python
recommended_k = recommended_kb // 1000 if recommended_kb > 0 else 8
max_peak_k = max_peak_kb // 1000 if max_peak_kb > 0 else 6
```
The `max(0, ...)` silently clamps an *actually-negative* budget to zero, and the `else 8` / `else 6` report fabricated 8k/6k numbers when the true budget is zero or unknown. Any downstream consumer — the dashboard, decompose.md, the future loop runner — reads 8k and proceeds as if it's safe.
**Fix**:
1. Drop the `max(0, ...)` clamp. Keep `net_kb` as the true arithmetic value (may be negative or zero). Already-floored callers (e.g. the dashboard) can compute `max(0, ...)` themselves; `vram_detect.py` returns the honest number.
2. Drop the `else 8` / `else 6` fallbacks. Report the real quotient even when zero.
3. Add a **warning line** to stdout when `net_kb <= 0` or `model_context_kb == 0` (see R3 for the loop refuse). For non-loop CLI invocations this is a human-readable warning, not an error exit.
```python
recommended_k = recommended_kb // 1000
max_peak_k = max_peak_kb // 1000
if recommended_kb <= 0:
print("WARNING: recommended context budget is zero or negative; "
"no usable context headroom for the configured system.")
```
### R3. Refuse unknown models (`model_context_kb: 0`) in loop mode
Today `detect_model_context` returns `0` on unknown models and `recommend_context` silently falls through to the VRAM/RAM branches. A loop tick with an unknown model could still proceed against an arbitrarily-deranged budget.
**Fix**: Add a `--loop-mode` flag to `vram_detect.py`'s CLI. When set:
- `model_context_kb == 0` is a hard error → print `"ERROR: model context window is unknown in --loop-mode. Set 'Override context window' in config.md or pass --model."` and exit `2`.
- `net_kb < 16000` is a hard error → print `"ERROR: available context ({}k) below 16k floor in --loop-mode (D13)."` and exit `2`.
The flag is optional. Non-loop callers (the dashboard, manual invocations) keep current behavior — only loops opt into the strict check. The future `loop-runner.py` will invoke `vram_detect.py --loop-mode --json` and expect either a 0 exit with a `{...}` JSON payload, or a 2 exit with a refuse message.
User's explicit `Override context window` in `config.md` (see `_parse_config_model`) is authoritative per D13: if a user has set an override, `detect_model_context` returns that override directly and the `model_context_kb == 0` refuse never fires. The flow already honors this — no special code needed.
### R4. Expose `available_context_kb` in JSON output
The loop runner needs a single authoritative figure for its per-tick budget. Today it would have to derive it from `recommended_kb - framework_overhead_tokens` itself, duplicating math.
**Fix**: Add `available_context_kb` to the JSON output block in `main()`:
```python
output = {
"gpu_vram_gb": gpu_vram_gb,
"ram_gb": ram_gb,
"model_context_kb": model_context_kb,
"framework_overhead_tokens": overhead_tokens,
"recommended_kb": recommended_kb, # net of overhead, before headroom
"recommended_k": recommended_k,
"headroom": headroom_pct / 100.0,
"max_peak_context_kb": max_peak_kb, # per-subtask peak (loop worktrees consume this)
"available_context_kb": max_peak_kb, # alias consumed by loop-runner.py; explicit field
"loop_mode_eligible": max_peak_kb >= 16000, # boolean: passes the 16k floor check
}
```
`available_context_kb` = `max_peak_kb` (post-R1 value, headroom applied exactly once). Two field names for the same number so both human-readable names and the runner's contract field are stable.
### R5. Add `## Loop Role Models` section to `config.md`
`config.md` today only documents VRAM settings. The loop system needs an explicit place for users to declare which model/session plays each role. Per `design/loops/functional.md` §5, three roles exist: `Implement:`, `Verify:`, `Orchestrate:`. The framework never inspects the *model* of each role (D8/D12) — it only needs to know which harness session to invoke per role, which is a harness-level concern that `loop.json` already handles via `roles.{implement,verify,orchestrate}.prompt`. So `config.md` should document the *expectation*, not encode it.
**Fix**: Append a new `## Loop Role Models` section to `~/.automaton/config.md`:
```markdown
## Loop Role Models
Loop ticks run three session roles. Roles are *sessions*, not models — a single model can fill multiple roles. Configure each loop's role-to-prompt binding in its `loop.json`; this section documents the framework's expectations only.
- **Implement:** — produces the artifact for this tick. Bound to `prompts/loop-implement.md` by default.
- **Verify:** — grades the artifact and emits the JSON verdict `{pass, score, reasons, next_hint}`. Bound to `prompts/loop-verifier.md`. The framework never inspects this role's model (D8); only its session.
- **Orchestrate:** — applies the verdict, calls exactly one `status.py` operation per tick, enforces brakes. Bound to `prompts/loop-orchestrate.md`.
Conflict-of-interest rule (D12): `Verify:` and `Implement:` must never be the same *session*. When two distinct sessions are infeasible (single-session harness), the runner falls back to session-only divergence — still safe.
Role context tiers are set per-loop in `loop.json`, not globally. The 16k floor (D13) applies regardless of tier.
```
### R6. Add 4k tier to `decompose.md` and tighten the table
`prompts/decompose.md` (around :82-84 per the design audit) has a context budget table that omits the small-context 4k tier that a single implement role might fit in when overhead + task brief is small. Per the design audit it also states the 16k floor.
**Fix**: Open `prompts/decompose.md`, find the existing context budget table (search for `4k` or `context` near the cited lines), add a row for the 4k tier and an explicit "≤ 16k: refuse" line above the table. Exact edits to be confirmed by reading the file at implementation time — this requirement locks the intent, not the diff.
If the existing table already covers 4k, this requirement is satisfied without edits; otherwise it is added. The 16k floor is the only hard refuse — 4k is a per-subtask peak recommendation, not a floor.
## Acceptance Criteria
- [ ] `recommend_context` returns `max_peak_kb` with headroom applied exactly once (verified by reading the function body — no inner `* (100 - headroom_pct) // 100` at the three budget-construction sites).
- [ ] `vram_detect.py`'s JSON output no longer reports 8 / 6 for `recommended_k` / `max_peak_k` when the underlying budget is zero. The actual quotients (including 0) are emitted.
- [ ] `vram_detect.py --loop-mode` exits `2` with the refuse message when the computed available context is `< 16000` tokens OR `model_context_kb == 0`.
- [ ] Without `--loop-mode`, the script preserves prior non-zero behavior on unknown / zero budgets (only a warning is added; no exit code change).
- [ ] JSON output includes `available_context_kb` and `loop_mode_eligible` fields.
- [ ] `config.md` includes the `## Loop Role Models` section verbatim (text may be condensed, intent preserved).
- [ ] `prompts/decompose.md` either acknowledges an existing 4k tier in its table or grows a 4k tier row, plus a `≤ 16k: refuse` line.
- [ ] New tests in `tests/` (Python) cover: double-headroom removed (regression test), `--loop-mode` refuse paths, JSON field presence, decompose.md tier presence.
- [ ] Pre-existing framework tests stay green: `python3 -m pytest tests/ -v`.
## Non-Goals
- Per-tick context budget enforcement **inside `vram_detect.py`** — that lives in `loop-runner.py` (task `add-loop-runner`). This task only *exposes* the numbers.
- Removing `model_context_kb == 0` fallback-to-VRAM in non-loop mode — that behavior is preserved for human CLI calls.
- Changing how `Override context window` is parsed (it's already authoritative per D13; this task just relies on it).
- Touching `loop-verifier.md` prompt contents — that's task `add-loop-templates-onboarding` (task 6). This task only adds a `## Loop Role Models` reference section to `config.md`.
## Dependencies
None. This is the first task in the bootstrap queue; downstream brakes and runner depend on it.
## Out of Scope (handled by Tier 2 `design/context-sizing/`)
Per D17: comprehensive context-sizing cleanup (`last-read-sha`, drift detection, decompose.md full rework, dashboard "model context" panel) is a sibling design driven by the first loop workstream after task 7 lands. This task limits itself to the six Tier 1 items above.