vault backup: 2026-06-09 23:15:17
This commit is contained in:
@@ -1,373 +1,296 @@
|
||||
# System Understanding Report — Loop / Scheduled Autonomy OOM
|
||||
|
||||
- **Flow id**: `recurring-bug-loop-oom` (pilot flow for autonomy ↔ deep-debug binding)
|
||||
- **Branch**: `fix/loop-scheduled-autonomy-oom`
|
||||
- **Worktree**: `E:\Source_code\Claude-code-bast-loop-scheduled-oom-fix`
|
||||
- **Author**: back-filled from existing working-tree diff (no commits ahead of `main`)
|
||||
- **Status**: `report` (this document) — pending human approval before `regression-test` advances
|
||||
|
||||
---
|
||||
tags:
|
||||
- OOM
|
||||
- 调度任务
|
||||
- 内存溢出
|
||||
- autonomy
|
||||
- 两阶段提交
|
||||
create time: 2026-06-09 22:30
|
||||
---
|
||||
|
||||
## 1. Problem
|
||||
# Loop / Scheduled Autonomy OOM 修复报告
|
||||
|
||||
### Symptom
|
||||
## 概述
|
||||
|
||||
Long-running sessions with active scheduled tasks (cron) and/or HEARTBEAT-driven proactive ticks accumulated growing memory, eventually OOM'ing the Bun process. The visible signature was:
|
||||
长时间运行的会话在活跃的定时任务(cron)和心跳驱动的主动循环下内存持续增长,最终导致 Bun 进程 OOM。根因是三个独立不足的缺陷在负载下交织:定时 tick 无同源去重、后台 fork 的 slash 命令提前报告成功、死进程记录永久阻塞去重。修复方案采用同源去重 + 进程印记 + 过期回收 + 延迟完成握手 + 两阶段提交排序。
|
||||
|
||||
- `runs.json` under `.claude/autonomy/` growing toward the 200-record cap with most entries stuck at `queued` or `running`
|
||||
- The internal command queue in REPL / headless mode draining slower than scheduled fires arrive
|
||||
- Each new fire calling `prepareAutonomyTurnPrompt`, which loads `AGENTS.md` + `HEARTBEAT.md` text and merges due-task lists into a fresh string, holding more closure state per pending command
|
||||
## 正文
|
||||
|
||||
### Expected behaviour
|
||||
### 基本信息
|
||||
|
||||
When a scheduled task fires while its prior run is still queued or running, the new fire should be **skipped** rather than enqueued behind it. When the process that started a run dies, the run should be reaped, not left as `running` forever. Background work spawned by a slash command should complete the originating autonomy run only when that background work itself finishes.
|
||||
- **Flow id**: `recurring-bug-loop-oom`(autonomy 与 deep-debug 绑定的先导 flow)
|
||||
- **分支**: `fix/loop-scheduled-autonomy-oom`
|
||||
- **状态**: `report`(本文档)——等待人工批准后推进到 `regression-test`
|
||||
|
||||
### Actual behaviour (before fix)
|
||||
### 问题现象
|
||||
|
||||
1. `useScheduledTasks` and the headless streaming path called `createAutonomyQueuedPrompt` unconditionally on every tick.
|
||||
2. `commitAutonomyQueuedPrompt` called `commitPreparedAutonomyTurn` *before* the run record was persisted, so even a duplicate fire that should have been dropped already mutated heartbeat-task last-run state.
|
||||
3. `AutonomyRunRecord` had no owner identity, so a run started by a now-dead process stayed `running` indefinitely. Subsequent runs of the same `sourceId` could not detect that their predecessor was effectively gone.
|
||||
4. Slash commands that forked detached background work (KAIROS / proactive paths) returned from `processUserInput` immediately. The harness in `handlePromptSubmit` then called `finalizeAutonomyRunCompleted`, marking the run `succeeded` while the actual work continued in the background — but the next scheduled tick of the same source could now race against that detached work, and any error in the detached work had no autonomy run to attribute to.
|
||||
#### 症状
|
||||
|
||||
### Reproduction shape
|
||||
长时间运行的会话在活跃的定时任务(cron)和/或 HEARTBEAT 驱动的主动循环下内存持续增长,最终 OOM 杀死 Bun 进程。可见特征:
|
||||
|
||||
Not a single deterministic repro — load-induced. Rough recipe:
|
||||
- `.claude/autonomy/` 下的 `runs.json` 趋向 200 条上限,大多数条目卡在 `queued` 或 `running`
|
||||
- REPL / headless 模式下的内部命令队列消耗速度慢于定时触发速度
|
||||
- 每次新触发都调用 `prepareAutonomyTurnPrompt`,加载 `AGENTS.md` + `HEARTBEAT.md` 文本并合并 due-task 列表到新字符串,每个 pending command 持有更多闭包状态
|
||||
|
||||
- Configure two `HEARTBEAT.md` tasks at `every 30s` interval
|
||||
- Add three cron tasks at `every 1m`
|
||||
- Let the session run > 1 hour, especially across a backgrounded slash command (e.g. KAIROS `/sleep`-style detached fork)
|
||||
- Watch `.claude/autonomy/runs.json` active-status entry count and Bun heap RSS
|
||||
#### 期望行为
|
||||
|
||||
### User impact
|
||||
当定时任务在先前运行仍在 `queued` 或 `running` 时触发,新触发应该被**跳过**而不是排队。当启动运行的进程死亡时,运行应该被回收,而不是永远留在 `running`。slash 命令生成的后台工作应该只在后台工作本身完成时才完成 originating autonomy run。
|
||||
|
||||
Sessions with long-lived autonomy/cron use cases were unsafe. The OOM took the entire CLI down, dropping any unflushed messages, MCP connections, and bridge state. Because `.claude/autonomy/` persists, restart did not heal — stale `running` records from the dead PID kept blocking dedup logic on the next start.
|
||||
#### 实际行为(修复前)
|
||||
|
||||
---
|
||||
1. `useScheduledTasks` 和 headless streaming 路径在每个 tick 上无条件调用 `createAutonomyQueuedPrompt`
|
||||
2. `commitAutonomyQueuedPrompt` 在 run record 持久化**之前**就调用了 `commitPreparedAutonomyTurn`,所以即使是应该被丢弃的重复触发也已经修改了心跳任务的 last-run 状态
|
||||
3. `AutonomyRunRecord` 没有 owner 标识,所以由已死进程启动的运行永远留在 `running`。后续同一 `sourceId` 的运行无法检测到其前身已经消失
|
||||
4. fork 了 detached 后台工作的 slash 命令(KAIROS / proactive 路径)立即从 `processUserInput` 返回。`handlePromptSubmit` 中的 harness 随后调用 `finalizeAutonomyRunCompleted`,将运行标记为 `succeeded`——但实际工作还在后台继续,同一 source 的下一个定时 tick 可能与该 detached 工作竞争
|
||||
|
||||
## 2. System boundary
|
||||
#### 复现方式
|
||||
|
||||
### In scope
|
||||
不是单一确定性复现——负载诱发。大致配方:
|
||||
|
||||
- Autonomy run lifecycle: create → running → succeeded / failed / cancelled (`src/utils/autonomyRuns.ts`)
|
||||
- Scheduled-task firing path: cron scheduler → REPL command queue (`src/hooks/useScheduledTasks.ts`)
|
||||
- Headless streaming variant of the same path (`src/cli/print.ts` `runHeadlessStreaming`)
|
||||
- Prompt-submit pipeline that finalizes runs after `processUserInput` returns (`src/utils/handlePromptSubmit.ts`)
|
||||
- Slash-command processing where a command may defer completion to background work (`src/utils/processUserInput/processUserInput.ts`, `processSlashCommand.tsx`)
|
||||
- `ToolUseContext` extension that lets non-bundled harnesses exercise the KAIROS-gated background-fork path (`src/Tool.ts`)
|
||||
- 配置两个 `HEARTBEAT.md` 任务,间隔 `every 30s`
|
||||
- 添加三个 cron 任务,间隔 `every 1m`
|
||||
- 让会话运行超过 1 小时,尤其跨过后台 slash 命令(如 KAIROS `/sleep` 风格的 detached fork)
|
||||
- 观察 `.claude/autonomy/runs.json` 活跃状态条目数和 Bun heap RSS
|
||||
|
||||
### Out of scope
|
||||
#### 用户影响
|
||||
|
||||
- The cron scheduler itself (`src/utils/cronScheduler.ts`) — its tick semantics are not changing
|
||||
- `autonomyFlows.ts` flow state machine — separate from per-run tracking
|
||||
- HEARTBEAT.md scheduling semantics — unchanged. `parseHeartbeatAuthorityTasks`
|
||||
does change narrowly by masking fenced code blocks before scanning so
|
||||
documented `tasks:` examples cannot shadow the real config block.
|
||||
- `prepareAutonomyTurnPrompt` content shape — only its call ordering relative to run creation changes
|
||||
- Any provider-level behaviour (`services/api/**`) — not touched
|
||||
> [!warning]
|
||||
> 长期运行 autonomy/cron 用例的会话不安全。OOM 会杀死整个 CLI,丢失未刷新的消息、MCP 连接和 bridge 状态。因为 `.claude/autonomy/` 持久化,重启无法治愈——死 PID 的 stale `running` 记录在下次启动时继续阻塞去重逻辑。
|
||||
|
||||
### Assumptions
|
||||
### 系统边界
|
||||
|
||||
- `process.pid` is stable for the lifetime of a Bun process and unique enough on a single host that a dead-PID heuristic is safe (collision risk acknowledged but bounded by `runs.json` retention).
|
||||
- `isProcessRunning(pid)` (from `genericProcessUtils.js`) returns `false` only when the process is actually gone; transient permission errors return `true`/safe-fail. Verified in step 6.
|
||||
- `getSessionId()` is initialized before any autonomy run creates records, since autonomy runs only originate after REPL or headless main loop boot.
|
||||
#### 范围内
|
||||
|
||||
---
|
||||
- Autonomy run 生命周期:create → running → succeeded / failed / cancelled(`src/utils/autonomyRuns.ts`)
|
||||
- 定时任务触发路径:cron scheduler → REPL command queue(`src/hooks/useScheduledTasks.ts`)
|
||||
- 同路径的 headless streaming 变体(`src/cli/print.ts` `runHeadlessStreaming`)
|
||||
- `processUserInput` 返回后 finalize runs 的 prompt-submit 管道(`src/utils/handlePromptSubmit.ts`)
|
||||
- 可能将完成延迟到后台工作的 slash 命令处理(`src/utils/processUserInput/processUserInput.ts`、`processSlashCommand.tsx`)
|
||||
- `ToolUseContext` 扩展,让非打包 harness 可以使用 KAIROS 门控的后台 fork 路径(`src/Tool.ts`)
|
||||
|
||||
## 3. Entry points
|
||||
#### 范围外
|
||||
|
||||
| Surface | Entry | Notes |
|
||||
- cron 调度器本身(`src/utils/cronScheduler.ts`)
|
||||
- `autonomyFlows.ts` flow 状态机
|
||||
- HEARTBEAT.md 调度语义
|
||||
- `prepareAutonomyTurnPrompt` 内容形状
|
||||
- 任何 provider 级行为
|
||||
|
||||
### 关键文件
|
||||
|
||||
| 文件 | 变更行数 | 重要性 |
|
||||
|---|---|---|
|
||||
| REPL | `useScheduledTasks` cron tick | Calls `createScheduledTaskQueuedCommand` (new helper) instead of raw `createAutonomyQueuedPrompt` |
|
||||
| REPL | Slash command pipeline | `processUserInput → processUserInputBase → processSlashCommand` now threads `autonomy` context so commands can defer completion |
|
||||
| Headless | `runHeadlessStreaming` cron path | Same migration to `createAutonomyQueuedPromptIfNoActiveSource`, plus `shouldCreate` callback honouring `inputClosed` |
|
||||
| Tool harness | `ToolUseContext.options.allowBackgroundForkedSlashCommands` | Non-prod way to exercise the KAIROS-gated detached-fork path; production still requires `feature('KAIROS')` + `AppState.kairosEnabled` |
|
||||
| Persistence | `.claude/autonomy/runs.json` | Schema gains `ownerProcessId`, `ownerSessionId`; readers must tolerate older records lacking these fields |
|
||||
| `src/utils/autonomyRuns.ts` | +260 | 拥有新的 identity + dedup + stale-recovery 逻辑;引入 `createAutonomyRunIfNoActiveSource`、`hasActiveAutonomyRunForSource`、`recoverStaleActiveAutonomyRun`、`commitAutonomyQueuedPromptIfNoActiveSource`、两阶段提交 |
|
||||
| `src/utils/processUserInput/processSlashCommand.tsx` | +707 / -454 | 重写 slash 命令派发,使 detached 后台工作可以 signal `deferAutonomyCompletion` |
|
||||
| `src/hooks/useScheduledTasks.ts` | +47 | 迁移两个 scheduler 调用点到 dedup helper |
|
||||
| `src/cli/print.ts` | +19 / -27 | headless 变体的相同迁移 |
|
||||
| `src/utils/handlePromptSubmit.ts` | +12 | 跟踪 `deferredAutonomyRunIds`,跳过 finalize |
|
||||
| `src/utils/processUserInput/processUserInput.ts` | +10 | 穿透 `autonomy` 上下文 |
|
||||
| `src/Tool.ts` | +6 | 添加 `allowBackgroundForkedSlashCommands` 测试逃生口 |
|
||||
|
||||
---
|
||||
### 调用流(修复后)
|
||||
|
||||
## 4. Key files
|
||||
#### 定时任务路径
|
||||
|
||||
| File | Lines changed | Why it matters |
|
||||
|---|---|---|
|
||||
| `src/utils/autonomyRuns.ts` | +260 | Owns the new identity + dedup + stale-recovery logic; introduces `createAutonomyRunIfNoActiveSource`, `hasActiveAutonomyRunForSource`, `recoverStaleActiveAutonomyRun`, `commitAutonomyQueuedPromptIfNoActiveSource`, two-phase commit. The structural heart of the fix. |
|
||||
| `src/utils/processUserInput/processSlashCommand.tsx` | +707 / -454 | Rewrites slash-command dispatch so detached background work signals `deferAutonomyCompletion`; refactor changes shape but not the public command set. |
|
||||
| `src/hooks/useScheduledTasks.ts` | +47 | Migrates both scheduler call sites to the dedup helper; extracts `createScheduledTaskQueuedCommand` for unit testing. |
|
||||
| `src/cli/print.ts` | +19 / -27 | Headless variant of the same migration; collapses the previous prepare+commit two-call sequence into the new dedup helper with `shouldCreate`. |
|
||||
| `src/utils/handlePromptSubmit.ts` | +12 | Tracks `deferredAutonomyRunIds` so it skips finalizing runs whose owning command deferred completion. |
|
||||
| `src/utils/processUserInput/processUserInput.ts` | +10 | Threads `autonomy` context and surfaces `deferAutonomyCompletion` on the result type. |
|
||||
| `src/Tool.ts` | +6 | Adds `allowBackgroundForkedSlashCommands` escape hatch for non-bundled harnesses (unit tests). |
|
||||
| `src/utils/__tests__/autonomyRuns.test.ts` | +168 | Regression coverage for dedup + stale recovery + ownership stamping. |
|
||||
| `src/hooks/__tests__/useScheduledTasks.test.ts` | new (75 lines) | Asserts scheduler does not double-fire while previous run is queued. |
|
||||
| `src/utils/processUserInput/__tests__/processSlashCommand.test.ts` | new (~280 lines) | Covers the deferred-completion handshake on slash-command paths. |
|
||||
|
||||
---
|
||||
|
||||
## 5. Call flow (post-fix)
|
||||
|
||||
```text
|
||||
cron tick (useScheduledTasks)
|
||||
└─> createScheduledTaskQueuedCommand(task)
|
||||
└─> createAutonomyQueuedPromptIfNoActiveSource
|
||||
├─> prepareAutonomyTurnPrompt (loads AGENTS.md + HEARTBEAT.md)
|
||||
├─> shouldCreate? ──► no ──► RETURN null (no side effects)
|
||||
└─> commitAutonomyQueuedPromptIfNoActiveSource
|
||||
└─> commitAutonomyQueuedPromptInternal(skipWhenActiveSource = true)
|
||||
└─> createAutonomyRunIfNoActiveSource
|
||||
├─> buildAutonomyRunRecord (stamps ownerProcessId, ownerSessionId)
|
||||
└─> persistAutonomyRunRecord(skip = true)
|
||||
└─> withAutonomyPersistenceLock
|
||||
├─> for each run with same (trigger,sourceId,ownerKey) and active status:
|
||||
│ ├─> isStaleActiveAutonomyRun? ──► recoverStaleActiveAutonomyRun (mark failed)
|
||||
│ └─> else ──► hasBlockingActiveRun = true
|
||||
├─> if blocking ──► RETURN created=false (no enqueue)
|
||||
└─> else ──► unshift record, write file, return true
|
||||
├─> if run is null ──► RETURN null (caller drops the tick)
|
||||
└─> else ──► commitPreparedAutonomyTurn(prepared) (heartbeat last-run state ONLY now mutates)
|
||||
└─> assemble QueuedCommand and return
|
||||
```mermaid
|
||||
graph TD
|
||||
A["cron tick useScheduledTasks"] --> B["createScheduledTaskQueuedCommand(task)"]
|
||||
B --> C["createAutonomyQueuedPromptIfNoActiveSource"]
|
||||
C --> D["prepareAutonomyTurnPrompt"]
|
||||
C --> E{"shouldCreate?"}
|
||||
E -->|否| F["RETURN null 无副作用"]
|
||||
E -->|是| G["commitAutonomyQueuedPromptIfNoActiveSource"]
|
||||
G --> H["commitAutonomyQueuedPromptInternal(skipWhenActiveSource=true)"]
|
||||
H --> I["createAutonomyRunIfNoActiveSource"]
|
||||
I --> J["buildAutonomyRunRecord 打印 ownerProcessId, ownerSessionId"]
|
||||
I --> K["persistAutonomyRunRecord(skip=true)"]
|
||||
K --> L{"withAutonomyPersistenceLock"}
|
||||
L --> M{"同 trigger+sourceId+ownerKey 的活跃运行?"}
|
||||
M -->|是-过期| N["recoverStaleActiveAutonomyRun 标记 failed"]
|
||||
M -->|是-未过期| O["hasBlockingActiveRun = true"]
|
||||
M -->|否| P["unshift record, write file"]
|
||||
O --> Q["RETURN created=false"]
|
||||
P --> R["commitPreparedAutonomyTurn 心跳状态才更新"]
|
||||
```
|
||||
|
||||
Two structural moves: (a) preparing the prompt no longer commits heartbeat state; only successful run insertion commits it. (b) blocking active runs of the same source short-circuit before the queue is touched.
|
||||
两个结构性改动:(a) 准备 prompt 不再提交心跳状态;只有成功插入 run 才提交。(b) 同源阻塞活跃运行在触及队列之前就短路。
|
||||
|
||||
For slash commands:
|
||||
#### Slash 命令路径
|
||||
|
||||
```text
|
||||
processUserInput → processUserInputBase
|
||||
└─> processSlashCommand(..., autonomy = cmd.autonomy)
|
||||
└─> command implementation
|
||||
├─> runs synchronously ──► returns normal result
|
||||
└─> spawns detached/background work ──► returns result with deferAutonomyCompletion = true
|
||||
+ handles its own finalize* call when work ends
|
||||
```mermaid
|
||||
graph TD
|
||||
A["processUserInput"] --> B["processUserInputBase"]
|
||||
B --> C["processSlashCommand(autonomy=cmd.autonomy)"]
|
||||
C --> D{"命令实现"}
|
||||
D -->|同步完成| E["返回正常结果"]
|
||||
D -->|生成 detached 后台工作| F["返回 result + deferAutonomyCompletion=true"]
|
||||
F --> G["自行处理 finalize 调用"]
|
||||
|
||||
handlePromptSubmit (caller of processUserInput):
|
||||
├─> records cmd.autonomy.runId in autonomyRunIds
|
||||
├─> on result with deferAutonomyCompletion=true: adds runId to deferredAutonomyRunIds
|
||||
└─> finalize loop: skips deferred ids in BOTH success and error branches
|
||||
H["handlePromptSubmit"] --> I["记录 cmd.autonomy.runId"]
|
||||
I --> J{"deferAutonomyCompletion=true?"}
|
||||
J -->|是| K["添加 runId 到 deferredAutonomyRunIds"]
|
||||
J -->|否| L["正常 finalize"]
|
||||
K --> M["finalize 循环: 跳过 deferred ids"]
|
||||
```
|
||||
|
||||
---
|
||||
### 数据流
|
||||
|
||||
## 6. Data flow
|
||||
|
||||
### `runs.json` record schema (delta)
|
||||
#### runs.json 记录 schema(增量)
|
||||
|
||||
```ts
|
||||
type AutonomyRunRecord = {
|
||||
// existing
|
||||
// 已有
|
||||
runId: string
|
||||
status: 'queued' | 'running' | 'succeeded' | 'failed' | 'cancelled'
|
||||
trigger: AutonomyTriggerKind
|
||||
sourceId?: string
|
||||
ownerKey?: string
|
||||
// new
|
||||
ownerProcessId?: number // process.pid at create time and at markRunning time
|
||||
ownerSessionId?: string // getSessionId() at the same points
|
||||
// ...
|
||||
// 新增
|
||||
ownerProcessId?: number // 创建时和 markRunning 时的 process.pid
|
||||
ownerSessionId?: string // 同一时机的 getSessionId()
|
||||
}
|
||||
```
|
||||
|
||||
Backward compatibility: older records with both fields absent are treated as "owner unknown" — they never satisfy `isStaleActiveAutonomyRun` (which requires `typeof ownerProcessId === 'number'`), so they remain blocking until they are completed normally or manually cancelled. This is intentional: we cannot prove they are stale.
|
||||
> [!info]
|
||||
> 向后兼容:两个字段都缺失的旧记录被视为"owner 未知"——它们永远不满足 `isStaleActiveAutonomyRun`(要求 `typeof ownerProcessId === 'number'`),所以保持阻塞直到正常完成或手动取消。这是有意的:我们无法证明它们是 stale 的。
|
||||
|
||||
### Stale-recovery rule
|
||||
#### 过期回收规则
|
||||
|
||||
```text
|
||||
isStaleActiveAutonomyRun(run) ⇔
|
||||
run.status ∈ {queued, running}
|
||||
∧ typeof run.ownerProcessId === 'number'
|
||||
∧ !isProcessRunning(run.ownerProcessId)
|
||||
isStaleActiveAutonomyRun(run) <=>
|
||||
run.status in {queued, running}
|
||||
&& typeof run.ownerProcessId === 'number'
|
||||
&& !isProcessRunning(run.ownerProcessId)
|
||||
```
|
||||
|
||||
Recovery mutates the in-memory list inside the persistence lock and writes it back, marking the stale run `failed` with error prefix `"Recovered stale active autonomy run"`.
|
||||
回收在持久化锁内修改内存列表并写回,将 stale run 标记为 `failed`,error 前缀为 `"Recovered stale active autonomy run"`。
|
||||
|
||||
### Heartbeat last-run state mutation point
|
||||
#### 心跳 last-run 状态变更点
|
||||
|
||||
Before fix: `commitAutonomyQueuedPrompt` called `commitPreparedAutonomyTurn(prepared)` *first*, then created the run. A skipped duplicate already advanced heartbeat last-run timestamps.
|
||||
- **修复前**:`commitAutonomyQueuedPrompt` **先**调用 `commitPreparedAutonomyTurn(prepared)`,然后创建 run。被跳过的重复触发已经推进了心跳 last-run 时间戳。
|
||||
- **修复后**:`commitPreparedAutonomyTurn` 只在 `createAutonomyRunIfNoActiveSource` 返回非 null 记录后才调用。被跳过的重复触发不影响心跳状态,所以下一个合格窗口仍在原始调度点。
|
||||
|
||||
After fix: `commitPreparedAutonomyTurn` is called only after `createAutonomyRunIfNoActiveSource` returns a non-null record. Skipped duplicates leave heartbeat state untouched, so the next eligible window is still at the originally scheduled point.
|
||||
### 状态模型
|
||||
|
||||
---
|
||||
#### Run 状态生命周期
|
||||
|
||||
## 7. State model
|
||||
|
||||
### Run status lifecycle (unchanged at edges, tightened in the middle)
|
||||
|
||||
```text
|
||||
queued ──► running ──► succeeded
|
||||
│ │
|
||||
│ └────► failed
|
||||
├──────────────────► cancelled
|
||||
└──► failed (stale recovery, new path)
|
||||
```mermaid
|
||||
graph TD
|
||||
A["queued"] --> B["running"]
|
||||
B --> C["succeeded"]
|
||||
B --> D["failed"]
|
||||
A --> E["cancelled"]
|
||||
A --> F["failed 过期回收新路径"]
|
||||
```
|
||||
|
||||
### New invariants
|
||||
#### 新不变量
|
||||
|
||||
1. **Same-source mutual exclusion**: at most one record with `(trigger, sourceId, ownerKey, status ∈ active)` is *non-stale* at any time. Enforced inside `withAutonomyPersistenceLock` in `persistAutonomyRunRecord`.
|
||||
1. **同源互斥**:任意时刻最多一条 `(trigger, sourceId, ownerKey, status in active)` 的非 stale 记录。在 `persistAutonomyRunRecord` 的 `withAutonomyPersistenceLock` 内强制执行。
|
||||
|
||||
2. **Owner stamping at active transitions**: any path that sets a run to `queued` or `running` must stamp `ownerProcessId = process.pid` and `ownerSessionId = getSessionId()`. `markAutonomyRunRunning` updated to do this for the running transition (creation already did it).
|
||||
2. **活跃转换时打 owner 印记**:任何将 run 设置为 `queued` 或 `running` 的路径都必须打印 `ownerProcessId = process.pid` 和 `ownerSessionId = getSessionId()`。`markAutonomyRunRunning` 已更新以在 running 转换时执行此操作。
|
||||
|
||||
3. **Two-phase commit ordering**: heartbeat-task last-run state may only be advanced after the run record has been successfully inserted. Equivalent to "prompt commit ⇒ run row exists".
|
||||
3. **两阶段提交排序**:心跳任务 last-run 状态只能在 run record 成功插入后才能推进。等价于"prompt commit => run row exists"。
|
||||
|
||||
4. **Deferred completion contract**: if a slash command's result has `deferAutonomyCompletion=true`, the harness (`handlePromptSubmit`) MUST NOT finalize the run; the command implementation OWNS the finalize call. Tracked via `deferredAutonomyRunIds` set scoped to a single `executeUserInput` invocation.
|
||||
4. **延迟完成契约**:如果 slash 命令的 result 带有 `deferAutonomyCompletion=true`,harness(`handlePromptSubmit`)**不得** finalize run;命令实现**拥有** finalize 调用。通过 `deferredAutonomyRunIds` 集合跟踪。
|
||||
|
||||
### Concurrency / retry risks
|
||||
#### 并发 / 重试风险
|
||||
|
||||
- Two processes sharing the same project root can race on `runs.json`. Mitigated by `withAutonomyPersistenceLock` (file-locking already in place), not by the new code.
|
||||
- Two ticks of the same scheduled task within a single process serialize on the same lock; only the first wins, the rest see the active record and return `null`.
|
||||
- A process killed between persisting the record and committing the prompt leaves a `queued` record with the dead PID. Stale recovery on the next tick of the same source converts it to `failed`, freeing the source. This is the new safety net.
|
||||
- 两个共享同一项目根目录的进程可以竞争 `runs.json`。由 `withAutonomyPersistenceLock`(文件锁)缓解。
|
||||
- 同一进程内同一定时任务的两个 tick 在同一把锁上串行;只有第一个获胜,其余看到活跃记录并返回 `null`。
|
||||
- 进程在持久化记录和提交 prompt 之间被杀死会留下带死 PID 的 `queued` 记录。同一 source 的下一个 tick 的过期回收将其转为 `failed`,释放 source。
|
||||
|
||||
### Two-phase commit crash window (acknowledged limitation)
|
||||
#### 两阶段提交崩溃窗口(已知限制)
|
||||
|
||||
Within `commitAutonomyQueuedPromptInternal` the order is:
|
||||
在 `commitAutonomyQueuedPromptInternal` 内,顺序是:
|
||||
|
||||
1. `createAutonomyRunCore` → `persistAutonomyRunRecord` → run row written under lock
|
||||
2. `commitPreparedAutonomyTurn(prepared)` → in-memory `heartbeatTaskLastRunByKey` Map advanced
|
||||
1. `createAutonomyRunCore` → `persistAutonomyRunRecord` → run row 在锁下写入
|
||||
2. `commitPreparedAutonomyTurn(prepared)` → 内存 `heartbeatTaskLastRunByKey` Map 推进
|
||||
|
||||
These two steps are NOT atomic. If the process is killed between (1) and (2):
|
||||
这两步**不是原子的**。如果进程在 (1) 和 (2) 之间被杀死:
|
||||
|
||||
- `runs.json` has a fresh `queued` record stamped with the now-dead PID.
|
||||
- `heartbeatTaskLastRunByKey` was an in-memory Map; its state vanishes with
|
||||
the process. On restart the Map is empty.
|
||||
- The dead-PID record is reaped via stale-recovery on the next tick of the
|
||||
same source → `status=failed`. New record can be created.
|
||||
- Because the Map starts empty after restart, every heartbeat task fires
|
||||
immediately on first tick rather than waiting for its configured
|
||||
interval window from the previous run.
|
||||
- `runs.json` 有一条带死 PID 的新鲜 `queued` 记录
|
||||
- `heartbeatTaskLastRunByKey` 是内存 Map;其状态随进程消失
|
||||
- 重启后 Map 为空,所有心跳任务在首次 tick 时立即触发
|
||||
|
||||
**Severity**: low. The Map is a runtime cache, not a persisted schedule
|
||||
contract; "fire immediately on restart" is a recoverable behaviour, not
|
||||
data corruption or duplicate work (the dead-PID record blocks the source
|
||||
until stale-recovery, so duplicate fires don't stack).
|
||||
> [!info]
|
||||
> **严重性**:低。Map 是运行时缓存,不是持久化调度契约;"重启后立即触发"是可恢复行为,不是数据损坏。死 PID 记录阻塞 source 直到过期回收,所以重复触发不会堆积。
|
||||
>
|
||||
> **为什么现在不修复**:在同一个锁内持久化心跳 last-run 状态会耦合两个不相关的状态机,成本超过罕见边界情况。已跟踪以便未来 flow 处理。
|
||||
|
||||
**Why not fix now**: persisting the heartbeat last-run state to disk inside
|
||||
the same lock would couple two unrelated state machines (autonomy runs vs
|
||||
heartbeat scheduling) and require a new on-disk schema. The cost outweighs
|
||||
the rare edge case (process death within microseconds between two
|
||||
in-memory operations). Tracked here so a future flow can pick it up if
|
||||
restart-after-crash schedule disruption becomes observable in practice.
|
||||
### 根因分析
|
||||
|
||||
---
|
||||
#### H1 — "Prompt 大小是 OOM 来源"
|
||||
|
||||
## 8. Existing tests
|
||||
**主张**:每个定时 tick 重建长 prompt 字符串;队列中这些字符串的累积保留导致堆压力。
|
||||
|
||||
### Pre-fix
|
||||
**支持证据**:`prepareAutonomyTurnPrompt` 确实每次 tick 构建多段字符串;`AGENTS.md` 有 220 行。
|
||||
|
||||
- `src/utils/__tests__/autonomyRuns.test.ts` covered create / list / mark transitions for the basic happy path.
|
||||
- No coverage for: dedup of same-source active run, stale-PID recovery, ownership stamping, deferred completion handshake, two-phase commit ordering.
|
||||
- `useScheduledTasks` had no unit tests — only indirect coverage via REPL integration.
|
||||
- `processSlashCommand` had no autonomy-context coverage.
|
||||
**反对证据**:diff 没有缩小任何 prompt 内容。如果 H1 是真正原因,修复应该把字符串组装放到缓存或 LRU 后面。
|
||||
|
||||
### Added in this branch
|
||||
**结论**:最多是贡献因素。作为主因被拒绝。
|
||||
|
||||
- `src/utils/__tests__/autonomyRuns.test.ts`: +168 lines covering dedup, stale recovery (mocked dead PID), ownership stamping at create + `markAutonomyRunRunning`, two-phase commit invariant.
|
||||
- `src/hooks/__tests__/useScheduledTasks.test.ts`: new file, 75 lines. Asserts scheduler skips double-fire when prior run is `queued`/`running`, and resumes when prior run finalizes.
|
||||
- `src/utils/processUserInput/__tests__/processSlashCommand.test.ts`: new file, ~280 lines. Covers `deferAutonomyCompletion=true` propagation; uses `allowBackgroundForkedSlashCommands` to bypass the `feature('KAIROS')` gate inside unit tests.
|
||||
#### H2 — "后台 fork 的 slash 命令泄漏 runs"
|
||||
|
||||
### Not yet covered (proposed for `regression-test` step)
|
||||
**主张**:KAIROS 风格的 slash 命令 fork detached 工作后立即返回;harness 随后将 run finalize 为 `succeeded`。后台工作中的任何错误都无法归属,且同一 source 的下一个定时触发发现没有活跃 run,多个后台 worker 在同一 source 后堆积。
|
||||
|
||||
- Cross-process race against the persistence lock — currently relies on file-lock correctness; consider a focused integration test that spawns two children and verifies only one wins.
|
||||
- Heartbeat last-run-state non-advance on skipped duplicates — assertable with a thin unit test against `prepareAutonomyTurnPrompt` + the dedup path; not blocking.
|
||||
**支持证据**:diff 显式添加了 `deferAutonomyCompletion`,将 `autonomy` 上下文穿透到 `processUserInputBase`,并更改 `handlePromptSubmit` 跳过延迟 run 的 finalize。
|
||||
|
||||
---
|
||||
**结论**:真实且承重。由针对性代码确认。
|
||||
|
||||
## 9. Competing root-cause hypotheses
|
||||
#### H3 — "定时任务 tick 对先前运行无去重"
|
||||
|
||||
### H1 — "Prompt size is the OOM source"
|
||||
**主张**:cron tick / heartbeat tick 无条件触发;如果先前 tick 的 run 仍在 `queued` / `running`,队列每个 interval 增长一条。跨多个 source 复合后,队列 + `runs.json` 活跃子集永不缩小。
|
||||
|
||||
**Claim**: each scheduled tick rebuilds a long prompt string (AGENTS.md + HEARTBEAT.md + due-task list); the cumulative retention of these strings in the queue causes heap pressure.
|
||||
**支持证据**:修复前 `useScheduledTasks` 和 `runHeadlessStreaming` 都调用 `createAutonomyQueuedPrompt`(无去重)。diff 用 `createAutonomyQueuedPromptIfNoActiveSource` 替换了两个调用点。
|
||||
|
||||
**Evidence for**: `prepareAutonomyTurnPrompt` does build a multi-section string each tick; `AGENTS.md` in this repo is now 220 lines.
|
||||
**结论**:真实且承重。由针对性代码确认。
|
||||
|
||||
**Evidence against**: the diff does not shrink any prompt content nor change `prepareAutonomyTurnPrompt`'s output. If H1 were the real cause, the fix would have moved string assembly behind a cache or LRU. The fix instead targets the *number* of in-flight runs.
|
||||
#### H4 — "死进程 run 永久毒化去重"
|
||||
|
||||
**Verdict**: contributing factor at most. Rejected as primary root cause.
|
||||
**主张**:即使 H3 修复了,进程在 run 期间被杀死会在磁盘上留下没有 owner 存活检查的 `running` 记录;下次加载 `runs.json` 的进程会将其视为阻塞,永远不再调度该 source。
|
||||
|
||||
### H2 — "Background-forked slash commands leak runs"
|
||||
**支持证据**:diff 打印 `ownerProcessId` 并添加 `isStaleActiveAutonomyRun` 检查。没有 H4,H3 的修复会创建新的失败模式(静默永久抑制)。
|
||||
|
||||
**Claim**: KAIROS-style slash commands that fork detached work return immediately from `processUserInput`; the harness in `handlePromptSubmit` then finalizes the run as `succeeded`. Any error in the background work is unattributable, and (more importantly) the *next* scheduled fire of the same source happens to find no active run, so multiple background workers stack up behind the same source.
|
||||
**结论**:真实但是次要的。它存在是因为 H3 的修复引入了它。必须一起发布。
|
||||
|
||||
**Evidence for**: the diff explicitly adds `deferAutonomyCompletion`, threads `autonomy` context into `processUserInputBase`, and changes `handlePromptSubmit` to skip finalization for deferred runs. New test file `processSlashCommand.test.ts` is dedicated to this exact handshake.
|
||||
> [!question] 为什么之前的本地补丁可能失败?
|
||||
> 这三个缺陷中的任何一个单独看起来都可以作为小 guard 修复,但只修复一个会将 OOM 转换为不同的错误行为(崩溃后静默抑制,或重复 detached worker)。最小正确修复需要所有三个原语:**同源去重**、**owner 印记 + 过期回收**、**延迟完成握手**,加上确保心跳状态在跳过的重复触发上永不推进的**两阶段提交排序**。
|
||||
|
||||
**Evidence against**: a pure same-source dedup miss would also explain the symptom; H3 covers that.
|
||||
### 修复计划
|
||||
|
||||
**Verdict**: real and load-bearing. Confirmed by the targeted code added.
|
||||
#### 最小修复面
|
||||
|
||||
### H3 — "Scheduled-task tick has no dedup against prior run"
|
||||
|
||||
**Claim**: cron tick / heartbeat tick fires unconditionally; if previous tick's run is still `queued`/`running` the queue grows by one each interval. Compounded across multiple sources, queue + `runs.json` active subset never shrink.
|
||||
|
||||
**Evidence for**: pre-fix `useScheduledTasks` and `runHeadlessStreaming` both called `createAutonomyQueuedPrompt` (no dedup). Diff replaces both call sites with `createAutonomyQueuedPromptIfNoActiveSource`. Persistence-side dedup added in the same change.
|
||||
|
||||
**Evidence against**: alone, this would make scheduling buggy but not necessarily OOM; the queue might catch up under light load.
|
||||
|
||||
**Verdict**: real and load-bearing. Confirmed by the targeted code added.
|
||||
|
||||
### H4 — "Dead-process runs poison dedup forever"
|
||||
|
||||
**Claim**: even with H3 fixed, a process killed mid-run leaves a `running` record on disk with no owner liveness check; the next process loading `runs.json` would treat it as blocking and never schedule that source again.
|
||||
|
||||
**Evidence for**: the diff stamps `ownerProcessId` and adds `isStaleActiveAutonomyRun` checked against `isProcessRunning`. Without H4, H3's fix would create a new failure mode (silent permanent suppression).
|
||||
|
||||
**Evidence against**: pre-fix code had no dedup, so this failure mode could not have been reached pre-fix.
|
||||
|
||||
**Verdict**: real, but secondary. It exists because H3's fix introduces it. Required to ship together.
|
||||
|
||||
---
|
||||
|
||||
## 10. Chosen root cause
|
||||
|
||||
**Combined H2 + H3 + H4**: the unbounded growth of active autonomy runs is the product of three independently insufficient gaps that line up under load:
|
||||
|
||||
1. Scheduled / heartbeat ticks do not dedup against an active prior run for the same source (H3).
|
||||
2. Background-forked slash commands report `succeeded` to the harness while their work is still detached, so subsequent ticks see no active run and stack workers behind the source (H2).
|
||||
3. Process death between record creation and run completion leaves zombie active records on disk that would block dedup permanently if (1) is fixed alone (H4).
|
||||
|
||||
Why previous local patches likely failed: any one of these in isolation looks fixable as a small guard, but fixing only one converts the OOM into a different misbehaviour (silent suppression after crash, or duplicate detached workers). The minimal correct fix needs all three primitives: **same-source dedup**, **owner stamping + stale recovery**, **deferred-completion handshake**, plus the **two-phase commit ordering** that ensures heartbeat state never advances on a skipped duplicate.
|
||||
|
||||
---
|
||||
|
||||
## 11. Fix plan
|
||||
|
||||
### Minimal fix surface
|
||||
|
||||
| Module | Change | Reason |
|
||||
| 模块 | 变更 | 原因 |
|
||||
|---|---|---|
|
||||
| `autonomyRuns.ts` | Owner stamping; `createAutonomyRunIfNoActiveSource`; `commitAutonomyQueuedPromptIfNoActiveSource`; two-phase commit; stale recovery | The structural primitives |
|
||||
| `useScheduledTasks.ts` | Replace both call sites with the dedup helper; extract `createScheduledTaskQueuedCommand` | Apply dedup at REPL scheduler |
|
||||
| `cli/print.ts` | Same migration in headless streaming path | Apply dedup in headless mode |
|
||||
| `handlePromptSubmit.ts` | Track `deferredAutonomyRunIds`; skip them in success and error finalize loops | Wire the deferred-completion contract |
|
||||
| `processUserInput.ts` | Thread `autonomy` ctx; surface `deferAutonomyCompletion` | Plumbing for the contract |
|
||||
| `processSlashCommand.tsx` | Background-fork commands set `deferAutonomyCompletion`; own their finalize call | Implementation of the contract |
|
||||
| `Tool.ts` | `allowBackgroundForkedSlashCommands` flag on `ToolUseContext.options` | Make the path testable from non-bundled harnesses |
|
||||
| `autonomyRuns.ts` | Owner 印记;`createAutonomyRunIfNoActiveSource`;`commitAutonomyQueuedPromptIfNoActiveSource`;两阶段提交;过期回收 | 结构性原语 |
|
||||
| `useScheduledTasks.ts` | 用 dedup helper 替换两个调用点 | 在 REPL scheduler 应用去重 |
|
||||
| `cli/print.ts` | headless streaming 路径的相同迁移 | 在 headless 模式应用去重 |
|
||||
| `handlePromptSubmit.ts` | 跟踪 `deferredAutonomyRunIds`;在 success 和 error finalize 循环中跳过它们 | 连接延迟完成契约 |
|
||||
| `processUserInput.ts` | 穿透 `autonomy` ctx;暴露 `deferAutonomyCompletion` | 契约的 plumbing |
|
||||
| `processSlashCommand.tsx` | 后台 fork 命令设置 `deferAutonomyCompletion`;拥有 finalize 调用 | 契约的实现 |
|
||||
| `Tool.ts` | `allowBackgroundForkedSlashCommands` 标志 | 使路径可从非打包 harness 测试 |
|
||||
|
||||
### Tests added
|
||||
#### 添加的测试
|
||||
|
||||
- `autonomyRuns.test.ts`: dedup, stale recovery (mocked dead PID via `isProcessRunning` mock), owner stamping at both create and `markAutonomyRunRunning`, two-phase commit ordering.
|
||||
- `useScheduledTasks.test.ts`: scheduler skips double-fire, resumes after finalize.
|
||||
- `processSlashCommand.test.ts`: deferred-completion handshake propagates to `handlePromptSubmit` correctly.
|
||||
- `autonomyRuns.test.ts`:去重、过期回收(mock 死 PID)、owner 印记、两阶段提交不变量
|
||||
- `useScheduledTasks.test.ts`:scheduler 跳过重复触发,finalize 后恢复
|
||||
- `processSlashCommand.test.ts`:延迟完成握手正确传播到 `handlePromptSubmit`
|
||||
|
||||
### Compatibility / migration risk
|
||||
#### 兼容性 / 迁移风险
|
||||
|
||||
- Older `runs.json` records lacking `ownerProcessId` are tolerated — never identified as stale, so they keep their blocking semantics. Operators who upgrade with stale `running` records on disk from a previous OOM crash will still need to manually `cancel` those runs (or wait for them to age out of the 200-record cap) the *first* time. After one full create cycle on the upgraded version, all new records carry owners.
|
||||
- **Observability gap on legacy blocking (added by reviewer 2026-04-28)**: when a no-owner active record blocks dedup, the current code path is silent — operators see "scheduled tasks stop firing" with no diagnostic. `implement` step MUST add a one-line warn log inside `persistAutonomyRunRecord`'s blocking branch: when `hasBlockingActiveRun = true` AND the blocking run has `ownerProcessId === undefined`, emit `[autonomyRuns] blocked by legacy un-owned active run <runId> (createdAt=<ts>); cancel manually if this is a stale upgrade artifact`. ≤ 10 lines of code, converts silent hang into a diagnosable signal. Do **not** change behavior — just observability.
|
||||
- `ToolUseContext.options.allowBackgroundForkedSlashCommands` is opt-in and defaults absent; production harness behaviour unchanged.
|
||||
- No on-disk schema version bump required.
|
||||
- 缺少 `ownerProcessId` 的旧 `runs.json` 记录被容忍——永远不被识别为 stale,保持阻塞语义。升级时磁盘上有 stale `running` 记录的运维人员仍需在**首次**手动 `cancel` 这些 run。
|
||||
- **遗留阻塞的可观察性缺口**:当无 owner 的活跃记录阻塞去重时,当前代码路径是静默的。`implement` 步骤**必须**在 `persistAutonomyRunRecord` 的阻塞分支添加一行 warn 日志。
|
||||
- 无 on-disk schema 版本升级。
|
||||
|
||||
### Rollback plan
|
||||
#### 回滚计划
|
||||
|
||||
- Revert the working tree to `main`'s versions of all 8 files. The `runs.json` schema additions are tolerated by older code (extra fields ignored).
|
||||
- If a stale record is preventing scheduling after rollback, manually edit `runs.json` (status → `cancelled`) or run `/autonomy flow cancel` for affected flows.
|
||||
- No dependency, no build flag, no settings-file change is needed for rollback.
|
||||
- 将工作树 revert 到 `main` 版本的所有 8 个文件。`runs.json` schema 增量被旧代码容忍(额外字段被忽略)。
|
||||
- 如果 stale record 在回滚后阻止调度,手动编辑 `runs.json`(status → `cancelled`)。
|
||||
- 无依赖、无构建标志、无 settings 文件更改。
|
||||
|
||||
### Out of scope (intentionally)
|
||||
### 验证
|
||||
|
||||
- Capping `prepareAutonomyTurnPrompt` output size (H1) — addressable later if needed; not load-bearing for the OOM.
|
||||
- Cross-process file-lock correctness review — relies on the existing `withAutonomyPersistenceLock`. Out of scope for this flow.
|
||||
- A migration utility to clean stale records on startup — discussed and rejected as avoidable: 200-record cap rolls them off naturally.
|
||||
|
||||
---
|
||||
|
||||
## 12. Verification
|
||||
|
||||
### Commands (binding per `.claude/autonomy/AGENTS.md` §4)
|
||||
#### 命令
|
||||
|
||||
```bash
|
||||
bun run typecheck
|
||||
@@ -379,114 +302,25 @@ bun run lint
|
||||
bun run build
|
||||
```
|
||||
|
||||
### Manual checks (proposed for `implement` step)
|
||||
#### 手动检查
|
||||
|
||||
- Start a session with two `HEARTBEAT.md` 30s tasks for ≥ 30 minutes; observe `runs.json` active-status entry count stays bounded (≤ number of distinct sources).
|
||||
- Force-kill the Bun process during a `running` record. Restart. Verify the next tick of the same source recovers (record marked `failed` with the stale-recovery error prefix) and a new run starts.
|
||||
- Run a KAIROS-gated detached slash command path under the test harness (`allowBackgroundForkedSlashCommands=true`) and verify `handlePromptSubmit` does not finalize the run while the background work is still active.
|
||||
- 启动带有两个 `HEARTBEAT.md` 30s 任务的会话,运行 30 分钟以上;观察 `runs.json` 活跃状态条目数保持有界
|
||||
- 在 `running` 记录期间强杀 Bun 进程。重启。验证同一 source 的下一个 tick 回收了记录(标记为 `failed`)并启动新 run
|
||||
- 在测试 harness 下运行 KAIROS 门控的 detached slash 命令路径,验证 `handlePromptSubmit` 在后台工作仍在活跃时不 finalize run
|
||||
|
||||
### Observability checks
|
||||
#### 可观察性检查
|
||||
|
||||
- `[ScheduledTasks] skipping <id>: previous run still queued or running` debug log appears when dedup fires (added in `useScheduledTasks.ts`). Use it to confirm dedup is reached in real sessions.
|
||||
- `runs.json` records with status `failed` and error starting `"Recovered stale active autonomy run"` indicate stale-recovery actually fired.
|
||||
- `[ScheduledTasks] skipping <id>: previous run still queued or running` debug 日志在去重触发时出现
|
||||
- `runs.json` 中 status `failed` 且 error 以 `"Recovered stale active autonomy run"` 开头的记录表明过期回收实际触发了
|
||||
|
||||
---
|
||||
### 未决问题
|
||||
|
||||
## 13. Open questions
|
||||
1. ~~`markAutonomyRunRunning` 是否在所有转换 autonomy run 到 `running` 的路径中被调用?~~ **已关闭(2026-04-28 验证)。** `markAutonomyRunRunning` 是**唯一**将 `AutonomyRunRecord.status` 转换为 `'running'` 的函数,无调用方绕过印记。
|
||||
|
||||
1. ~~Should `markAutonomyRunRunning` be called in *all* paths that transition an autonomy run to `running`, or only the prompt-submit path?~~ **Closed (verified 2026-04-28).**
|
||||
`markAutonomyRunRunning` (`autonomyRuns.ts:554-579`) is the **only** function that transitions `AutonomyRunRecord.status → 'running'`. It stamps `ownerProcessId = process.pid` and `ownerSessionId = getSessionId()` unconditionally, then internally calls `markManagedAutonomyFlowStepRunning` to mirror to flow state. `markManagedAutonomyFlowStepRunning` is only invoked from this one call site (`autonomyRuns.ts:571`); no caller bypasses the stamp. All four real callers (`cli/print.ts:2177`, `screens/REPL.tsx:4859`, `utils/handlePromptSubmit.ts:492`, `utils/swarm/inProcessRunner.ts:741`) go through the stamping path. Flow records intentionally do not carry owner fields — the run record is source of truth and flow steps mirror via `latestRunId`. Stale-recovery operates on runs, so flow-step runs are covered.
|
||||
2. ~~`getSessionId()` import was added to `autonomyRuns.ts`. Confirm no circular import is introduced...~~ **Closed (verified 2026-04-28).**
|
||||
No risk on three counts: (a) `autonomyRuns.ts:4` already imported `getProjectRoot` from `bootstrap/state.js`; the new `getSessionId` is appended to the same import line, adding zero new module-level coupling. (b) Reverse direction is empty — `grep -rn 'autonomy*' src/bootstrap/` yields no results, so the dependency stays one-way. (c) `getSessionId()` (`bootstrap/state.ts:425-427`) returns `STATE.sessionId`, which is initialized at module load with `randomUUID()` and re-randomized by `resetStateForTests()` per test — never `undefined`, never throws. The existing test file deliberately uses the real `bootstrap/state` module (not a mock) and already asserts `ownerProcessId === process.pid` / `ownerSessionId` is a string in the new ownership tests, plus exercises stale recovery with a fake dead PID (`2_147_483_647`). No mock updates needed.
|
||||
3. Is the 200-record cap still appropriate now that recovery turns stale runs into `failed`? Active records will churn faster; the cap may roll off legitimate completed records sooner. Not a correctness issue, but worth noting.
|
||||
2. ~~`getSessionId()` 导入是否引入循环依赖?~~ **已关闭(2026-04-28 验证)。** 无风险:反向依赖为空,`getSessionId()` 永不 `undefined`,永不抛出。
|
||||
|
||||
---
|
||||
3. 200 条上限在过期回收将 stale run 转为 `failed` 后是否仍然合适?活跃记录会更快轮转;上限可能更早滚掉合法完成记录。不是正确性问题,但值得记录。
|
||||
|
||||
## 14. Approval gate
|
||||
## 关联笔记
|
||||
|
||||
This SUR satisfies `AGENTS.md` §3 step `report` exit criteria once a human reviewer:
|
||||
|
||||
- [x] confirms the chosen root cause (§10) matches their reading of the diff — **agent-ticked under user delegation 2026-04-28; see §15 verification table row 1**
|
||||
- [x] approves the §11 fix plan including the deferred-completion contract — **agent-ticked under user delegation 2026-04-28; Concern A's warn-log requirement folded into §11**
|
||||
- [x] acknowledges the §11 compatibility note about pre-existing stale records on disk — **agent-ticked under user delegation 2026-04-28; §11 extended with Concern A observability gap**
|
||||
- [x] §13 open question 1 (stamping completeness in flow-step runners) — closed 2026-04-28; see §13 for the verification trace
|
||||
- [x] Concern B (processSlashCommand.tsx >50% diff) — **resolved 2026-04-28 by commit-split rule, see §15**
|
||||
|
||||
---
|
||||
|
||||
## 15. Reviewer findings (2026-04-28, agent-reviewed)
|
||||
|
||||
The user explicitly delegated SUR review work to the agent. The four §14 checkboxes
|
||||
remain user's decision; this section records the agent's verification work and
|
||||
recommendations to make that decision faster and more auditable.
|
||||
|
||||
### Verification work performed
|
||||
|
||||
| Claim | Cross-check | Result |
|
||||
|---|---|---|
|
||||
| §10 H2/H3/H4 互锁 | Walked each "fix only one" counterfactual | ✅ Real interlock — fixing only one converts OOM into a different bug (silent suppression / persistent stacking) |
|
||||
| §11 fix surface covers all 8 modified files | Compared against `git diff --stat` | ✅ Each file has a row in the table |
|
||||
| §11 "extra fields ignored" rollback claim | JSON parse semantics | ✅ Correct |
|
||||
| §11 compatibility claim "tolerated" | Re-read `isStaleActiveAutonomyRun` (`autonomyRuns.ts`) | ⚠️ Tolerance is real but **silent** — gap surfaced as Concern A below |
|
||||
| §13 Q1 owner stamping completeness | (closed in earlier turn — see §13) | ✅ |
|
||||
| §13 Q2 circular-import / mock impact | (closed in earlier turn — see §13) | ✅ |
|
||||
| §13 Q3 200-record cap acceptability | Reasoned about stale-recovery-driven churn | ✅ Non-blocking; forensic loss only |
|
||||
|
||||
### Concerns surfaced
|
||||
|
||||
**Concern A — silent legacy blocking (now folded into §11)**: when a no-owner active
|
||||
record from a pre-upgrade crash blocks dedup, the operator gets no signal — just
|
||||
"scheduled tasks stop firing." The §11 compatibility section was extended to require
|
||||
a one-line warn log in `implement`. This is an observability fix, not a behavior
|
||||
change.
|
||||
|
||||
**Concern B — `processSlashCommand.tsx` is +707/-454 (>50% rewrite)** — **RESOLVED 2026-04-28**:
|
||||
investigation showed the diff is composed of:
|
||||
- **18 contract-related lines** (verified by `grep -E '(autonomy|QueuedCommand|deferAutonomy|finalizeAutonomy|allowBackgroundForkedSlashCommands|deferredAutonomy)'`):
|
||||
- import `QueuedCommand` type
|
||||
- import `finalizeAutonomyRunCompleted` / `finalizeAutonomyRunFailed`
|
||||
- add `autonomy?: QueuedCommand['autonomy']` parameter to `executeForkedSlashCommand` (3 sites)
|
||||
- extend KAIROS gate to also accept `context.options.allowBackgroundForkedSlashCommands === true` (test escape hatch)
|
||||
- finalize the run from the detached background path on success/failure
|
||||
- set `deferAutonomyCompletion: Boolean(autonomy?.runId)` on the result
|
||||
- thread `autonomy` to nested calls
|
||||
- **~30-50 lines** of necessary control-flow scaffolding around the contract code
|
||||
- **~250 lines** of pure Biome reformatting churn (single-line imports, trailing semicolons)
|
||||
|
||||
**Resolution rule (binding for `implement`)**: when committing this branch, split
|
||||
`processSlashCommand.tsx` into **two commits** on the same branch:
|
||||
|
||||
```text
|
||||
chore: reformat processSlashCommand with Biome # ~250 lines, formatter-only
|
||||
feat: thread autonomy run id through forked slash commands for deferred completion # ~50 lines, contract logic
|
||||
```
|
||||
|
||||
This satisfies `~/.claude/rules/deep-debug/core.md` §2 ("bug fix 不允许混入...格式化")
|
||||
in spirit by making the contract commit reviewable in isolation, without
|
||||
requiring a fragile manual revert of formatter output (which Biome would
|
||||
re-apply on the next save). All other 7 modified files in the OOM fix do not
|
||||
require commit splitting — verify by sampling their diffs at `implement` time.
|
||||
|
||||
**Concern C — stale-recovery rate metric (deferred)**: post-implement, track daily
|
||||
stale-recovery count. If consistently elevated, the 200-record cap may need
|
||||
revisiting (relates to §13 Q3). Not a blocker; suggested for follow-up flow.
|
||||
|
||||
### Agent recommendations on the §14 checkboxes
|
||||
|
||||
| §14 box | Agent recommendation | Rationale |
|
||||
|---|---|---|
|
||||
| §10 chosen root cause | Approve | H2/H3/H4 互锁 verified; diff supports each branch |
|
||||
| §11 fix plan (with §15 Concern A folded in) | Approve | Minimal, complete, regression-tested |
|
||||
| §11 compatibility note | Acknowledge as-extended (§11 now includes the warn-log requirement from Concern A) | Silent legacy blocking would surprise users; the added log makes it diagnosable |
|
||||
| Concern B `processSlashCommand.tsx` >50% diff | Resolved by commit-split rule (chore + feat) | 18 lines contract + ~250 lines formatter churn; commit split makes review tractable without fragile revert |
|
||||
|
||||
**Final status (2026-04-28, agent-resolved under user delegation)**: all five §14
|
||||
boxes ticked. Flow `recurring-bug-loop-oom` may advance from `report` to
|
||||
`regression-test`. Implement-time obligations folded in:
|
||||
|
||||
1. Add the legacy-blocking warn log in `persistAutonomyRunRecord` (Concern A, ≤10 lines)
|
||||
2. Commit-split `processSlashCommand.tsx` into chore + feat (Concern B)
|
||||
3. Verify the other 7 modified files do not need commit-splitting (sample their diffs)
|
||||
4. Track stale-recovery counts post-deploy for §13 Q3 / Concern C follow-up
|
||||
|
||||
After approval: flow advances to `regression-test`. The targeted commands in §12 must produce a verifiable failing state on the *pre-fix* tree before the post-fix tree is allowed to satisfy `implement`. Since this branch already contains the fix, the regression evidence will be reconstructed by checking out one parent, running the targeted tests (expected: fail), then returning to HEAD (expected: pass).
|
||||
- [[sur-skill-overflow-bugs]]
|
||||
|
||||
Reference in New Issue
Block a user