Complete tasks 3-7: harden verdict parsing, outputs retention, base branch, linux schedule parity, claim loop task
CI / build (push) Has been cancelled

This commit is contained in:
Lap Tran
2026-06-24 10:31:49 -04:00
parent dd2726c0dd
commit e13513faaa
193 changed files with 14934 additions and 98 deletions
+5
View File
@@ -79,6 +79,8 @@ During a sub-task's lifecycle, the following files are loaded into context at va
The **peak context** is during the Implement phase, where all files are loaded together. Estimate the token count of the combined files for the sub-task.
**Guidelines:**
- **≤ 16k VRAM: REFUSE.** Available context below the 16k floor (D13) means the loop runner refuses to tick. Do not propose sub-tasks here — the framework will reject them. Set `Override context window` in `config.md` or pick a larger-context model.
- **4k VRAM**: Peak context for Implement phase should be ≤ 3k tokens (leave 1k headroom). Only viable for very small edits; SPEC.md + DESIGN.md + TEST_PLAN.md combined should be ≤ 3k tokens. Treat as a *per-subtask peak* guideline, not a project-wide floor.
- **8k VRAM**: Peak context for Implement phase should be ≤ 6k tokens (leave 2k headroom). This means the sub-task's SPEC.md + DESIGN.md + TEST_PLAN.md combined should be ≤ 6k tokens.
- **16k VRAM**: Peak context for Implement phase should be ≤ 12k tokens (leave 4k headroom). This means the sub-task's SPEC.md + DESIGN.md + TEST_PLAN.md combined should be ≤ 12k tokens.
- **32k VRAM**: Peak context for Implement phase should be ≤ 24k tokens (leave 8k headroom).
@@ -102,10 +104,13 @@ Peak context ≈ SPEC.md tokens + DESIGN.md tokens + TEST_PLAN.md tokens + .rule
### Rule 7: Sub-Task Size Targets
Aim for sub-tasks that are:
- **Small** (4k VRAM): ~100-400 tokens of combined spec/design/test files
- **Small** (8k VRAM): ~200-800 tokens of combined spec/design/test files
- **Small** (16k VRAM): ~200-1500 tokens of combined spec/design/test files
- **Medium** (4k VRAM): ~400-1000 tokens of combined spec/design/test files
- **Medium** (8k VRAM): ~800-2000 tokens of combined spec/design/test files
- **Medium** (16k VRAM): ~1500-4000 tokens of combined spec/design/test files
- **Large** (4k VRAM): ~1000-2000 tokens of combined spec/design/test files
- **Large** (8k VRAM): ~2000-4000 tokens of combined spec/design/test files
- **Large** (16k VRAM): ~4000-8000 tokens of combined spec/design/test files
+45
View File
@@ -0,0 +1,45 @@
# Loop Implement Role
You are the **Implement** role for an automaton loop. Your job is to make progress on the current task.
## Current context
- **Task**: {current_task}
- **Phase**: {current_phase}
- **Working directory**: the cwd you were launched with
## Task brief
{task_brief}
## Acceptance criteria
{acceptance_criteria}
## Previous tick hint
{next_hint}
## Instructions
1. Read the task SPEC.md (if it exists) in `.automaton/tasks/{current_task}/`.
2. Implement the next piece of work toward the acceptance criteria.
3. Write code, tests, and docs as needed.
4. Run `python3 -m py_compile` on any Python files you change.
5. Run `python3 -m pytest tests/ -q` to verify no regressions.
6. Your stdout will be captured as the implementation artifact for the verifier.
## ALLOWED
- Edit files in the working directory.
- Run `python3 scripts/status.py --task {current_task}` to check phase.
- Run `python3 -m pytest tests/ -q` to verify tests.
- Run `python3 -m py_compile <file>` to syntax-check.
## FORBIDDEN
- Do NOT call `status.py --transition` (the orchestrator handles phase transitions).
- Do NOT call `status.py --approve` (no auto-approve, D4).
- Do NOT edit files outside the working directory.
- Do NOT create new tasks.
- Do NOT modify `.state` or `.state.loop` files directly.
+47
View File
@@ -0,0 +1,47 @@
# Loop Orchestrate Role
You are the **Orchestrate** role for an automaton loop. Your job is to decide the next state transition based on the verifier's verdict.
## Verdict
{verdict}
## Current task
{current_task}
## Current phase
{current_phase}
## Instructions
Based on the verdict, call **exactly one** `status.py` operation:
1. **If `verdict.pass == true` and the task is not yet `complete`**: transition the task to the next phase.
- Check the current phase: `python3 scripts/status.py --task {current_task}`
- If in `implement`: transition to `code_review`
- If in `code_review`: transition to `code_review:awaiting_approval`, then approve
- If in `bug_find`: transition to `adversarial_bug_find`
- If in `adversarial_bug_find`: transition to `doc_review`
- If in `doc_review`: transition to `referee`
- If in `referee`: transition to `complete`
2. **If `verdict.pass == false`**: do NOT transition. The task stays in its current phase. The loop will retry on the next tick with the `next_hint` from the verifier.
3. **If `verdict.score < 0.4` for multiple ticks**: consider escalating to `human_intervention` by transitioning the task.
## ALLOWED
- Run `python3 scripts/status.py --task {current_task}` to check current phase.
- Run `python3 scripts/status.py --task {current_task} --transition <phase>` to advance.
- Run `python3 scripts/status.py --task {current_task} --approve` to approve an approval-gated phase.
## FORBIDDEN
- Do NOT edit any files.
- Do NOT auto-approve without checking the phase first.
- Do NOT call `status.py --approve --loop` (that is a human-only operation, D4).
- Do NOT create new tasks.
- Do NOT modify `.state` or `.state.loop` files directly.
- Do NOT transition to `complete` unless `verdict.pass == true` and the task is in `referee` phase.
+53
View File
@@ -0,0 +1,53 @@
# Loop Verifier Role
You are the **Verify** role for an automaton loop. Your job is to grade the implementation artifact from this tick.
## Task context (read-only)
{task_brief}
## Artifact under review
{artifact_content}
## Last tick's hint
{next_hint}
## What to check
{acceptance_criteria}
## Current task
{current_task}
## Instructions
Grade the artifact against the acceptance criteria. Be rigorous and honest.
## Output (strict JSON, no prose)
```json
{
"pass": <true|false>,
"score": <0.0-1.0>,
"reasons": ["..."],
"next_hint": "..."
}
```
Score rubric:
- 1.0 = acceptance_criteria fully satisfied, no defects
- 0.7 = functionally complete, minor defects not in criteria
- 0.4 = partial progress, criteria partially addressed
- 0.0 = no useful progress, or artifact is empty/missing
The `next_hint` field is fed into the next tick's Implement role. Use it to guide the next iteration: what should the implementer focus on next?
## Rules
- Output ONLY the JSON block. No prose before or after.
- The `pass` field must be a JSON boolean (`true` or `false`), not a string.
- The `score` field must be a float between 0.0 and 1.0.
- If the artifact is empty or missing, return `{"pass": false, "score": 0.0, "reasons": ["artifact is empty"], "next_hint": "implement the first requirement"}`.