## Execution Protocol — Planning Layer 🔵 > This is the execution engine. Every agent turn MUST follow this protocol. > It replaces reactive execution with structured planning. > No new code. No new agents. Pure protocol enforcement via this prompt. ### 0. Complexity Classification (MUST run first) Before ANY action, classify the task: **Simple → Direct Execution** | Criteria (ANY ONE) | Example | |--------------------|---------| | Single file read/write | "Read schema.ts" | | Single command | "npm test" | | Pure query | "What time is it?" | | Conversation | "Hello" | **Complex → MUST Use Task Tree** | Criteria (ANY ONE) | Example | |--------------------|---------| | 3+ files to modify | "Fix all FK references" | | Multi-step workflow | "Generate project → install → build → test" | | Schema/database change | "Add inspection_records table" | | Cross-project operation | "Fix all 3 generated projects" | | Auth/security change | "Add RBAC" | | New feature | "Add file upload" | | Bug root-cause analysis | "Why does FK constraint fail?" | | 2+ tool categories (read+write+exec) | "Read schema, fix it, rebuild" | | Estimated > 45 min or > 2 systems | "Epic-level task" | **Decision rule: if uncertain → treat as Complex.** --- ### 1. Task Tree Construction (Complex Tasks ONLY — MANDATORY) > ⚠️ Complex tasks MUST build a Task Tree BEFORE entering §2 PLAN. > No flat task lists. Task Tree is the only valid entry to Planning. #### 1a. Task Tree Definition Task Tree is an **execution tree**, not a feature tree. **Wrong (module tree):** ``` Project ├── User ├── Order └── Payment ``` **Correct (execution tree):** ``` Project ├── T1 Architecture Setup ├── T2 Database Schema [depends_on: T1] ├── T3 Backend API [depends_on: T2] ├── T4 Frontend [depends_on: T3] └── T5 Verification [depends_on: T4] ``` #### 1b. Mandatory Output Format Before §2 PLAN, output: ``` ### Task Tree | Task ID | Name | Depends On | Parallelizable | Est. Size | |---------|------|------------|----------------|-----------| | T1 | Project Bootstrap | NONE | NO | Small | | T2 | Database Schema | T1 | NO | Medium | | T3 | Auth Service | T2 | YES | Medium | | T4 | User Service | T2 | YES | Medium | | T5 | API Gateway | T3,T4 | NO | Medium | ``` Each Task ID maps to a subtask in §2 PLAN. #### 1c. Dependency Rules **Hard Dependency** — upstream task MUST finish before starting: ``` T3 Backend API depends_on: T2 Database Schema ``` **Soft Dependency** — can start in parallel, must integrate before final: ``` T4 UI Pages depends_on: Design Tokens (soft) ``` #### 1d. Parallelism Optimization Agent MUST identify parallelizable tasks. Output: ``` ### Parallel Groups **Group A (can run simultaneously):** - T3 Auth Service - T4 User Service **Group B (can run simultaneously):** - T6 Unit Tests - T7 Integration Tests ``` If all tasks form a serial chain (T1→T2→T3→...), **must explain why** parallelism is impossible. #### 1e. Learning Context Integration For each task, inject relevant lessons from the Learning Loop: ``` ### Per-Task Injected Risks **T3: Backend Generator** - [R1] AP-3 DDL Constraint Leakage — add isConstraintLine() before column parsing Source: reflections/2026-06-05-ddl-leakage.md - [R2] FK REPLACE → use ON CONFLICT DO UPDATE Source: best-practices.md #11 ``` #### 1f. Task Tree Quality Gate Before proceeding to §2 PLAN, verify ALL: - [ ] Root Task exists (exactly one task with depends_on: NONE) - [ ] No orphan tasks (every non-root task has a valid parent chain to root) - [ ] No circular dependency (A→B→C→A) - [ ] No XL Task (if estimated > 4h, split into sub-tasks) - [ ] No unjustified serial chain (if T1→T2→T3→T4 with no hard deps, explain or parallelize) - [ ] Every task has Injected Risks populated (§1e) - [ ] Agent can independently complete every task (no external human dependency) **If ANY gate fails: do NOT proceed to §2 PLAN. Re-split the tree.** --- ### 2. Planning Protocol (Complex Tasks ONLY) Before any tool call on a complex task: ``` ## PLAN ### Goal ### Subtasks 1. → Verify: 2. → Verify: ... ### Risks - - ### Verification Strategy - Build: - Tests: - Lint: ``` Then use `update_plan` if >3 subtasks. Then execute. ### 3. Pre-flight Check (Complex Tasks ONLY) Before executing step 1, **MUST** complete ALL of the following: #### 3a. Dependency Scan - [ ] Which files/tables/modules are involved? - [ ] What are their dependencies (FK, imports, types)? - [ ] What's the blast radius if I get it wrong? **Anti-pattern:** "I'll fix this FK error" → "Oh wait, here's another FK error" → "And another..." **Correct pattern:** "Let me scan ALL FK references first, then fix them all at once." #### 3b. Learning Context Injector (MANDATORY — 3 searches) **Search 1: Similar failures** ``` ctx_search(queries: [" failure pattern lesson"], source: "reflections", sort: "timeline") ``` **Search 2: Relevant best practices** ``` memory_search(query: " best practice checklist pattern") ``` **Search 3: Relevant playbook** ``` Check playbooks/ for matching task type ``` Inject top findings into PLAN > Risks. Format: ``` ### Injected Risks (from Learning Context) - [R1] — source: - [R2] — source: best-practices.md #N - [R3] — source: playbooks/.md ``` #### 3c. Domain Checklist Injection **Based on task type, append the relevant checklist to PLAN:** **SQL / Schema tasks:** - □ FK 引用:所有 FK 用 ON CONFLICT DO UPDATE,不用 REPLACE(SQLite REPLACE = DELETE+INSERT) - □ DDL 解析:PRIMARY KEY / CONSTRAINT 行不能当数据字段 - □ 迁移安全:先检查所有引用方,再改 schema - □ WAL 模式:多进程读写不阻塞 **JavaScript / TypeScript tasks:** - □ Falsy 陷阱:0、''、false 用 `??` 替代 `||`(|| 会把 0 判为 falsy) - □ 类型生成:entity names 复数→单数(Albums→Album),但不碰 status/bus 等词根 - □ 模板语法:模板字面量内嵌反引号用 \` 转义或字符串拼接 - □ 启发式检测:必须 2+ 信号同时匹配,单信号不可靠 **Next.js / Frontend tasks:** - □ 路径:Next.js App Router 用 `app/` 不是 `src/app/` - □ Tailwind content 路径匹配实际目录 - □ tsconfig paths 匹配实际目录结构 - □ JSX:三元表达式嵌套使用括号包裹,避免 `{}` 嵌套错误 **Backend / API tasks:** - □ generateTypes 必须包含 User 实体 → CreateUserInput 类型 - □ 所有 resource types 生成完整的 CRUD Input 类型 - □ route params 命名与 service 函数签名一致 **Generator / Builder tasks:** - □ 输出路径:验证生成目录结构与实际 runtime 预期一致 - □ DDL→TS:schema.ts 中 CONSTRAINT/FK 行不能被解析为 interface 字段 - □ Type mapping:SQL types 到 TS types 的映射表完整 - □ Template safety:模板字面量不输出裸反引号 **Release / CI tasks:** - □ Exit code 规范:PASS/WARN→0, FAIL→1 - □ 回退路径:每个 destructive 操作有对应回退步骤 - □ backward compat:新代码不改老接口 If no checklist matches, add: - □ No domain-specific checklist available — proceed with extra caution ### 4. Failure Replanning Protocol When a step fails: **DO NOT** fix the immediate error and continue blindly. **DO:** 1. Log: what failed, error message, which step 2. Analyze: is this a local fix or does it invalidate the plan? 3. If local → fix, verify, continue 4. If structural → **REPLAN** — output a revised PLAN section 5. If same step fails 3 times → **PAUSE** and ask user 6. **Auto-Capture Check** (after fix, before continuing): → See complete trigger rules in §6 Auto-Capture Trigger (T1-T7) → If ANY trigger matches → write structured Reflection to `reflections/YYYY-MM-DD-.md` → Reflection MUST contain ALL 7 required fields (Context/Root Cause/Fix/Pattern Class/Future Trigger/Checklist/Recurrence Count) → Skip only when: fix < 5 min + first occurrence + no test/build/lint failure 7. **Promotion Check** (after Reflection written): | Appearances | Action | Destination | |------------|--------|------------| | 1st | Reflection | reflections/ | | 2nd | Pattern | patterns/design-patterns.md (new section) | | 3rd | Best Practice | patterns/best-practices.md (numbered entry) | | 4th | Checklist Item | AGENTS.md §3c Domain Checklists | | 5th | Hard Rule | AGENTS.md §3c (bolded, marked ⚠️ HARD RULE) | ### 5. Verification Protocol (Complex Tasks ONLY) After all subtasks complete: ``` ## VERIFICATION - Build: [PASS/FAIL] - Tests: [PASS/FAIL / N suites / N tests] - Lint: [PASS/FAIL] - Regression risk: [NONE/LOW/MEDIUM/HIGH] - Manual changes needed: [NONE / ] ``` **MANDATORY GATE: Pre-flight Compliance Check** Before closing any complex task, verify ALL of: - [ ] PLAN contains `### Injected Risks (from Learning Context)` section - [ ] At least 2 knowledge sources were searched (ctx_search / memory_search / playbooks) - [ ] Matching domain checklist from §3c is included in PLAN - [ ] If failure occurred: structured Reflection was written to `reflections/YYYY-MM-DD-.md` - [ ] Reflection contains ALL 7 required fields (Context/Root Cause/Fix/Pattern Class/Future Trigger/Checklist/Recurrence) If ANY gate fails: **task is NOT complete**. Go back and fill the gap. ### 6. Governance Compliance 🔒 Compliance is mandatory for all complex task execution. - **State machine, heartbeat, blocker, replan, scope detection, health score, auto-stop, recovery mode** → see GOVERNANCE.md - **Multi-project portfolio management** → see PORTFOLIO.md - **MANDATORY Governance Quality Gate** must run before project output (GOVERNANCE.md §9) ### 7. Auto-Capture Trigger — COMPLETE Rule Set (after ANY failure) **ALWAYS check these conditions after fixing a failure. Write Reflection when ANY condition matches:** | # | Trigger Condition | Why | |---|------------------|-----| | T1 | `build` command failed | Build failure always indicates systemic error | | T2 | `test` command returned non-zero or unexpected failures | Test failure = regression risk | | T3 | `lint` / `tsc --noEmit` failed | Syntax/type errors in generated code | | T4 | Fix time > 30 minutes elapsed | Long fix = complex root cause worth documenting | | T5 | Same error class appeared before (check ctx_search / Promotion Log) | Recurrence → must upgrade | | T6 | Generated output path/format/structure mismatched target convention | E.g. `src/app/` vs `app/`, wrong SQL dialect, missing types | | T7 | Generated content caused runtime crash or data corruption | Silent data corruption is highest severity | **Skip reflection (not worth recording) when:** - Fix < 5 min AND never seen before AND no test/build/lint failure - Typo fix (single character, no logic change) - Expected external service failure (network timeout, API rate limit) **Reflection write checklist (ALL 7 fields REQUIRED):** ``` [ ] Context — what task was being performed [ ] Root Cause — the bug, not the symptom (banned: "以后注意", "下次小心") [ ] Fix — exact change, with code diff if applicable [ ] Pattern Class — abstract category (e.g. "SQLite FK", "JS Falsy Coercion", "Path Convention") [ ] Future Trigger — when to inject this lesson (e.g. "any SQL schema change") [ ] Checklist — ≥2 actionable checks for next time [ ] Recurrence Count — 1 (update on each recurrence) ``` --- ## Learning Loop 🔄 > Knowledge stored in reflections/patterns/playbooks MUST flow back into Planning. > This section closes the loop. ### Loop Architecture ``` Failure/Success ↓ Reflection (structured, with Pattern Class) ↓ Promotion Check (1→2→3→4→5 appearances) ↓ Pattern → Best Practice → Checklist → Hard Rule ↓ Pre-flight Injector (§3b) + Domain Checklists (§3c) ↓ Next Plan includes past lessons ↓ Execution ↓ Verification → triggers new Reflection if needed ↓ (loop closes) ``` ### Promotion Rules | Count | Promotion To | Location | |-------|-------------|----------| | 1st | Reflection | `reflections/YYYY-MM-DD-.md` | | 2nd | Design Pattern | `patterns/design-patterns.md` (new Anti-Pattern section) | | 3rd | Best Practice | `patterns/best-practices.md` (numbered entry) | | 4th | Domain Checklist | AGENTS.md §3c (update matching checklist, or create new) | | 5th | Hard Rule | AGENTS.md §3c (bolded, marked ⚠️ HARD RULE) | ### Quality Gate for Reflections A valid Reflection MUST include ALL of: - [ ] Context (what task) - [ ] Root Cause (the bug, not the symptom) - [ ] Fix (exact change) - [ ] Pattern Class (abstract category) - [ ] Future Trigger (when to re-inject) - [ ] Checklist (actionable items) - [ ] Recurrence Count **Rejected:** "以后注意", "下次小心", "当时没想清楚", "应该多测试" **Accepted:** "SQLite FK: REPLACE=DELETE+INSERT, use ON CONFLICT DO UPDATE" ### Storage Locations ``` reflections/ ← First occurrence of a failure pattern patterns/design-patterns ← Abstract reusable anti-patterns (≥2 occurrences) patterns/best-practices ← Validated rules (≥3 occurrences) AGENTS.md §3c ← Domain Checklists (≥4 occurrences) AGENTS.md §3c ⚠️ ← Hard Rules (≥5 occurrences) playbooks/ ← Successful multi-step workflows (not failure-driven) vault.md §Self-Evolution ← History of promotions (timeline) ``` --- ## Agent Package Context Budget ⚠️ HARD RULE > Established 2026-06-05. Every Agent Package MUST declare its token budget before execution. > This prevents context overflow, KV cache thrashing, and mid-task truncation in multi-agent workflows. ### Size → Token Budget Mapping | Level | Reading | Modifying | New Code | Dependencies | Token Budget | |-------|---------|-----------|----------|-------------|-------------| | Small | ≤3 files | ≤2 files | ≤300 lines | ≤1 | ≤30K tokens | | Medium | ≤8 files | ≤5 files | ≤800 lines | ≤3 | ≤80K tokens | | Large | ≤15 files | ≤10 files | ≤1500 lines | ≤5 | ≤200K tokens | | XL | ❌ FORBIDDEN | ❌ | ❌ | ❌ | ❌ | ### Input Size Pre-check (MANDATORY) Before dispatching any Agent Package, pre-compute: 1. **File inventory** — list every file the agent will read 2. **Token estimate** — for code: `bytes / 3`; for Chinese text: `chars / 1.5`; round up 10% for safety 3. **Budget check** — `sum(estimated_input_tokens) + prompt_overhead(2K) + output_buffer(20% of budget)` ≤ Token Budget ``` ### Context Budget Declaration **AP-N**: - Files to read: N (total ~X lines / ~Y estimated tokens) - Prompt overhead: ~2K tokens - Output buffer: ~Z tokens (20% of budget) - Total estimated: ~W tokens - Budget: V tokens - Utilization: (W/V) × 100% - ⚠️ OVER BUDGET: → if YES, split further ``` ### Over-budget Resolution If estimated utilization > 90%: 1. Split the package into sub-packages 2. Pre-index large reference files instead of reading them inline (use ctx_index → ctx_search) 3. Provide only relevant snippets, not entire files 4. Re-check budget after split ### Token Estimation Constants | Content Type | Token Ratio | Notes | |-------------|------------|-------| | TypeScript/JS | 1 token ≈ 3.5 chars | conservative estimate | | Python | 1 token ≈ 3.5 chars | | | SQL | 1 token ≈ 3.0 chars | shorter tokens | | Chinese text | 1 token ≈ 1.5 chars | CJK characters are token-heavy | | English prose | 1 token ≈ 4.0 chars | | | JSON/YAML | 1 token ≈ 3.0 chars | punctuation-heavy |