vault backup: 2026-06-09 23:15:17
This commit is contained in:
@@ -1,82 +1,96 @@
|
||||
# System Understanding Report — Skill Search / Skill Learning Overflow Bugs
|
||||
|
||||
- **Flow id**: `recurring-bug-skill-overflow` (sibling pilot to `recurring-bug-loop-oom`)
|
||||
- **Branch**: `fix/loop-scheduled-autonomy-oom` (folded into the OOM PR — same audit-and-cap pattern)
|
||||
- **Trigger**: post-merge review of the autonomy OOM fix surfaced unbounded module-level state in adjacent `EXPERIMENTAL_SKILL_SEARCH` and `SKILL_LEARNING` subsystems. The user explicitly asked for a `肯定也有同类溢出` audit.
|
||||
|
||||
---
|
||||
tags:
|
||||
- OOM
|
||||
- 内存溢出
|
||||
- 缓存治理
|
||||
create time: 2026-06-09 22:30
|
||||
---
|
||||
|
||||
## 1. Problem
|
||||
# Skill Search / Skill Learning 溢出 Bug 修复报告
|
||||
|
||||
The autonomy OOM bug came from unbounded module-level state (run records, scheduler queues, heartbeat timestamps) growing for the lifetime of the process. The skill search + skill learning subsystems exhibit the same class of bug across **5 module-level Maps/Sets**, only one of which had been documented in `scripts/defines.ts` ("projectContext cache 无淘汰机制(非 GB 级主因)").
|
||||
## 概述
|
||||
|
||||
These bugs were latent because:
|
||||
审计发现 `EXPERIMENTAL_SKILL_SEARCH` 和 `SKILL_LEARNING` 子系统存在 5 个无界模块级状态,与 autonomy OOM 属于同类 bug。修复方案采用 FIFO/LRU 淘汰 + 运行时双层门控,默认关闭实验特性直到运维人员显式启用。
|
||||
|
||||
- `EXPERIMENTAL_SKILL_SEARCH` / `SKILL_LEARNING` were enabled-by-default in `DEFAULT_BUILD_FEATURES`, but tests pass because they exercise short paths.
|
||||
- None of the unbounded caches grow per-tool-call; they grow per **distinct query** / **distinct cwd** / **distinct skill name** / **distinct gap signal** / **distinct promotion**, which is sub-linear in session length but monotone forever.
|
||||
- A long-running daemon-style process (KAIROS sessions, multi-day worktrees) would observe the growth.
|
||||
## 正文
|
||||
|
||||
## 2. Module-level state audit
|
||||
### 问题背景
|
||||
|
||||
| File:Line | Symbol | Pre-fix bound | Pre-fix evict |
|
||||
- **Flow id**: `recurring-bug-skill-overflow`(与 `recurring-bug-loop-oom` 为姊妹审计)
|
||||
- **分支**: `fix/loop-scheduled-autonomy-oom`(已合并到 OOM PR,使用相同的 audit-and-cap 模式)
|
||||
- **触发**: 合并 autonomy OOM 修复后,发现相邻的 `EXPERIMENTAL_SKILL_SEARCH` 和 `SKILL_LEARNING` 子系统存在同类无界模块级状态。用户明确要求进行 `肯定也有同类溢出` 审计。
|
||||
|
||||
### 问题分析
|
||||
|
||||
autonomy OOM 来自无界模块级状态(运行记录、调度队列、心跳时间戳)随进程生命周期不断增长。skill search + skill learning 子系统在 **5 个模块级 Map/Set** 上表现出相同类型的 bug,其中只有一个在 `scripts/defines.ts` 中有文档记录("projectContext cache 无淘汰机制(非 GB 级主因)")。
|
||||
|
||||
这些 bug 之前是潜伏的,原因如下:
|
||||
|
||||
- `EXPERIMENTAL_SKILL_SEARCH` / `SKILL_LEARNING` 在 `DEFAULT_BUILD_FEATURES` 中默认启用,但测试通过是因为它们只走短路径
|
||||
- 这些无界缓存不是按每次工具调用增长,而是按**不同查询 / 不同 cwd / 不同 skill 名称 / 不同 gap 信号 / 不同 promotion** 增长,相对于会话长度是亚线性的,但永远单调递增
|
||||
- 长时间运行的守护进程式进程(KAIROS 会话、多日 worktree)会观察到增长
|
||||
|
||||
### 模块级状态审计
|
||||
|
||||
| 文件:行 | 符号 | 修复前上限 | 修复前淘汰 |
|
||||
|---|---|---|---|
|
||||
| `intentNormalize.ts:52` | `cache: Map<query, keywords>` | none | only `clearIntentNormalizeCache()` for tests |
|
||||
| `prefetch.ts:17` | `discoveredThisSession: Set<skillName>` | none | none |
|
||||
| `prefetch.ts:18` | `recordedGapSignals: Set<gapKey>` | none | none |
|
||||
| `projectContext.ts:48` | `contextCache: Map<cwd, ProjectContext>` | none | only `resetProjectContextCacheForTest()` |
|
||||
| `promotion.ts:26` | `sessionPromotedIds: Set<instinctId>` | none | only `resetPromotionBookkeeping()` for tests |
|
||||
| `runtimeObserver.ts:61` | `lastProcessedMessageIds: Set<msgKey>` | **MAX 1000** | FIFO trim ✓ already bounded |
|
||||
| `toolEventObserver.ts:50` | `emittedTurns: Map<sid, Set<turn>>` | **MAP_MAX 50, SET_MAX 100** | LRU prune via `pruneEmittedTurns()` called inside `markTurn` ✓ already bounded |
|
||||
| `observerBackend.ts:21` | `registry: Map<name, Backend>` | fixed N | n/a — registry pattern, finite ✓ |
|
||||
| `intentNormalize.ts:52` | `cache: Map<query, keywords>` | 无 | 仅 `clearIntentNormalizeCache()` 测试用 |
|
||||
| `prefetch.ts:17` | `discoveredThisSession: Set<skillName>` | 无 | 无 |
|
||||
| `prefetch.ts:18` | `recordedGapSignals: Set<gapKey>` | 无 | 无 |
|
||||
| `projectContext.ts:48` | `contextCache: Map<cwd, ProjectContext>` | 无 | 仅 `resetProjectContextCacheForTest()` 测试用 |
|
||||
| `promotion.ts:26` | `sessionPromotedIds: Set<instinctId>` | 无 | 仅 `resetPromotionBookkeeping()` 测试用 |
|
||||
| `runtimeObserver.ts:61` | `lastProcessedMessageIds: Set<msgKey>` | **MAX 1000** | FIFO trim 已有上限 |
|
||||
| `toolEventObserver.ts:50` | `emittedTurns: Map<sid, Set<turn>>` | **MAP_MAX 50, SET_MAX 100** | LRU prune 已有上限 |
|
||||
| `observerBackend.ts:21` | `registry: Map<name, Backend>` | 固定 N | 不适用——注册表模式,有限 |
|
||||
|
||||
**5 unbounded out of 8 module-level mutables.** All 5 are addressed in this PR.
|
||||
> [!warning]
|
||||
> 8 个模块级可变状态中有 **5 个无界**。所有 5 个都在本次 PR 中修复。
|
||||
|
||||
## 3. Severity rationale
|
||||
### 严重性分析
|
||||
|
||||
Per-entry cost is small (key strings + small objects), so OOM in days is unlikely on a normal workstation. But the canary scenarios:
|
||||
每个条目的成本很小(key 字符串 + 小对象),所以正常工作站上几天内 OOM 不太可能。但以下场景值得关注:
|
||||
|
||||
- **`intentNormalize.cache`**: every distinct Chinese query → Haiku call → cached. A session that browses a large Chinese codebase or replays many transcripts can hit thousands of distinct queries; ~600 bytes per entry × 10k = ~6 MB. Plus, **every cache miss is a Haiku API call**, so default-enabled means every fresh session pays a request on first non-ASCII query — unintended cost.
|
||||
- **`projectContext.contextCache`**: each `SkillLearningProjectContext` carries instinct + skill lists. Multi-worktree orchestrators (this very repo!) blow past the typical "1 cwd per session" assumption.
|
||||
- **`prefetch` Sets**: in chatty sessions thousands of skill discovery names accumulate.
|
||||
- **`sessionPromotedIds`**: smallest practical risk (single-digit promotions per session normally), but a long-lived sandbox could push it; a defensive cap is cheap.
|
||||
- **`intentNormalize.cache`**:每个不同的中文查询 → Haiku 调用 → 缓存。浏览大型中文代码库或重放大量 transcript 的会话可以命中数千个不同查询;约 600 字节/条 x 10k = 约 6 MB。此外,**每次缓存未命中都是一次 Haiku API 调用**,默认启用意味着每个新会话在首次非 ASCII 查询时都要付费——这是意外成本。
|
||||
- **`projectContext.contextCache`**:每个 `SkillLearningProjectContext` 携带 instinct + skill 列表。多 worktree 编排器(就是这个仓库!)会突破典型的"每会话 1 个 cwd"假设。
|
||||
- **`prefetch` Sets**:在多话会话中,数千个 skill 发现名称会累积。
|
||||
- **`sessionPromotedIds`**:实际风险最小(单会话通常个位数 promotion),但长期运行的沙盒可能推高它;防御性上限成本很低。
|
||||
|
||||
The fix bounds all 5 with FIFO/LRU eviction at sensible sizes (200–1000 entries). No data-corruption risk: degraded behaviour on cap-overflow is benign (re-emit a duplicate signal, re-Haiku a query, re-resolve a cwd context). Same risk profile as the autonomy stale-recovery design.
|
||||
### 修复方案
|
||||
|
||||
## 4. Fix surface
|
||||
|
||||
| File | Change |
|
||||
| 文件 | 变更 |
|
||||
|---|---|
|
||||
| `src/services/skillSearch/intentNormalize.ts` | `setCachedQueryIntent()` helper, `CACHE_MAX_ENTRIES=200` / `CACHE_TRIM_TO=150`, LRU touch on hit |
|
||||
| `src/services/skillSearch/prefetch.ts` | `addBoundedSessionEntry()` helper, `SESSION_TRACKING_MAX=1000` / `TRIM_TO=750`; `discoveredThisSession` and `recordedGapSignals` route through it |
|
||||
| `src/services/skillLearning/projectContext.ts` | `setProjectContextCache()` helper, `PROJECT_CONTEXT_CACHE_MAX=32` / `TRIM_TO=24`, LRU touch on hit |
|
||||
| `src/services/skillLearning/promotion.ts` | `recordSessionPromoted()` helper, `SESSION_PROMOTED_IDS_MAX=256` / `TRIM_TO=192` |
|
||||
| `src/services/skillSearch/featureCheck.ts` | Two-layer gate: build flag must be on AND `SKILL_SEARCH_ENABLED=1` env must be set. Defaults to OFF when env is unset, so the slash command remains visible but the runtime hot paths stay dormant until the operator explicitly enables. |
|
||||
| `src/services/skillLearning/featureCheck.ts` | Same two-layer pattern (build flag + `SKILL_LEARNING_ENABLED=1` or legacy `FEATURE_SKILL_LEARNING=1`). |
|
||||
| `scripts/defines.ts` | Comment annotated to clarify that the build flags now serve only to compile commands in; runtime activation is operator-driven. |
|
||||
| `src/services/skillSearch/intentNormalize.ts` | `setCachedQueryIntent()` 辅助函数,`CACHE_MAX_ENTRIES=200` / `CACHE_TRIM_TO=150`,命中时 LRU touch |
|
||||
| `src/services/skillSearch/prefetch.ts` | `addBoundedSessionEntry()` 辅助函数,`SESSION_TRACKING_MAX=1000` / `TRIM_TO=750`;`discoveredThisSession` 和 `recordedGapSignals` 通过它路由 |
|
||||
| `src/services/skillLearning/projectContext.ts` | `setProjectContextCache()` 辅助函数,`PROJECT_CONTEXT_CACHE_MAX=32` / `TRIM_TO=24`,命中时 LRU touch |
|
||||
| `src/services/skillLearning/promotion.ts` | `recordSessionPromoted()` 辅助函数,`SESSION_PROMOTED_IDS_MAX=256` / `TRIM_TO=192` |
|
||||
| `src/services/skillSearch/featureCheck.ts` | 双层门控:构建标志必须开启 **且** `SKILL_SEARCH_ENABLED=1` 环境变量必须设置。环境变量未设置时默认关闭 |
|
||||
| `src/services/skillLearning/featureCheck.ts` | 同样的双层模式(构建标志 + `SKILL_LEARNING_ENABLED=1` 或 `FEATURE_SKILL_LEARNING=1`) |
|
||||
| `scripts/defines.ts` | 注释标注,说明构建标志现在仅用于编译命令;运行时激活由运维驱动 |
|
||||
|
||||
## 5. Why default-off (without removing from build)?
|
||||
修复用 FIFO/LRU 淘汰在合理大小(200-1000 条)上限制所有 5 个。没有数据损坏风险:达到上限时的行为是良性的(重新发出重复信号、重新 Haiku 查询、重新解析 cwd 上下文)。
|
||||
|
||||
Three reasons aside from the unbounded-cache concern:
|
||||
### 为什么默认关闭(而不是从构建中移除)
|
||||
|
||||
1. **Implicit cost**: `intentNormalize` calls Haiku on cache miss. Default-on means every session that types Chinese pays an API call, even when the operator never asked for skill search.
|
||||
2. **Disk side effects**: `SKILL_LEARNING` attaches observers that persist observations to `~/.claude` storage. Storage volume should be opt-in, not background.
|
||||
3. **Experimental status**: the flag is literally named `EXPERIMENTAL_*`. Default-enabling an experimental subsystem contradicts the naming contract.
|
||||
除了无界缓存问题外,还有三个原因:
|
||||
|
||||
**The fix is NOT to remove the flags from `DEFAULT_BUILD_FEATURES`** — doing so would also strip the `/skill-search` and `/skill-learning` slash commands from the build, leaving operators with no UI to opt in. Instead the activation logic in `featureCheck.ts` was changed to a two-layer gate:
|
||||
1. **隐式成本**:`intentNormalize` 在缓存未命中时调用 Haiku。默认开启意味着输入中文的每个会话都要付费一次 API 调用,即使运维人员从未要求 skill search。
|
||||
2. **磁盘副作用**:`SKILL_LEARNING` 附加观察者,将观察结果持久化到 `~/.claude` 存储。存储量应该是 opt-in,而不是后台行为。
|
||||
3. **实验状态**:标志名称就是 `EXPERIMENTAL_*`。默认启用实验子系统与命名约定矛盾。
|
||||
|
||||
- **Layer 1 (compile-time)**: `feature('EXPERIMENTAL_SKILL_SEARCH')` / `feature('SKILL_LEARNING')` must be on. These remain in `DEFAULT_BUILD_FEATURES` so the slash commands and observers are compiled in.
|
||||
- **Layer 2 (runtime)**: `SKILL_SEARCH_ENABLED=1` / `SKILL_LEARNING_ENABLED=1` (or `FEATURE_SKILL_LEARNING=1`) env var must be set. Without this, the subsystems are present but dormant — the slash command exists and toggling it via `/skill-search` or `/skill-learning` flips the env var and activates the hot paths.
|
||||
> [!info]
|
||||
> 修复**不是**从 `DEFAULT_BUILD_FEATURES` 中移除标志——这样做也会从构建中剥离 `/skill-search` 和 `/skill-learning` slash 命令,让运维人员没有 UI 来 opt-in。相反,`featureCheck.ts` 中的激活逻辑改为双层门控:
|
||||
>
|
||||
> - **第 1 层(编译时)**:`feature('EXPERIMENTAL_SKILL_SEARCH')` / `feature('SKILL_LEARNING')` 必须开启。这些保留在 `DEFAULT_BUILD_FEATURES` 中,以便 slash 命令和观察者被编译进来。
|
||||
> - **第 2 层(运行时)**:`SKILL_SEARCH_ENABLED=1` / `SKILL_LEARNING_ENABLED=1`(或 `FEATURE_SKILL_LEARNING=1`)环境变量必须设置。不设置的话,子系统存在但休眠——slash 命令存在,通过 `/skill-search` 或 `/skill-learning` 切换会翻转环境变量并激活热路径。
|
||||
>
|
||||
> 最终结果:运维人员在 UI 中看到切换开关,但子系统**关闭直到他们手动切换**。
|
||||
|
||||
Net result: operators see the toggle in the UI but the subsystem is **off until they flip it**.
|
||||
### 范围外(已提交 follow-up)
|
||||
|
||||
## 6. Out of scope (filed for follow-up)
|
||||
- **CI 测试失败**(`prefetch.test.ts`、`skillLearningSmoke.test.ts`)出现在本分支 CI 中。两个测试都**显式通过环境变量启用了特性**,所以默认禁用不是原因。它们是实验代码路径中的既有功能问题,值得单独 flow 处理。
|
||||
- **持久化层上限**(观察文件、instinct 注册表):`observationStore.ts` 已有 30 天清除和 1MB 归档阈值;`skillGapStore.ts` 使用有限状态生命周期。磁盘侧状态有适当限制;OOM 类问题严格在进程内状态。
|
||||
|
||||
- **Test failures on CI** (`prefetch.test.ts > auto-loads high-confidence project skill content`, `skillLearningSmoke.test.ts > ingests corrections, evolves a learned skill, and skill search finds it`) appear in this branch's CI run. Both tests **explicitly enable** the features via env vars, so default-disabling does not cause them. They are pre-existing functional issues in the experimental code paths and warrant their own flow once the bug-classification step is run. Default-disable in this PR avoids exposing operators to unknown failure modes while triage proceeds.
|
||||
- **Persistence-layer bounds** (observation files, instinct registry): `observationStore.ts` already has 30-day purge and 1MB archive thresholds; `skillGapStore.ts` uses a finite-state lifecycle. Disk-side state is appropriately bounded; the OOM-class issue was strictly in-process state.
|
||||
|
||||
## 7. Verification
|
||||
|
||||
Local checks (full suite covers cap behaviour via existing tests; the caps degrade gracefully so no test should break):
|
||||
### 验证
|
||||
|
||||
```bash
|
||||
bun run typecheck # 0 errors
|
||||
@@ -88,4 +102,8 @@ bun run lint
|
||||
bun run build
|
||||
```
|
||||
|
||||
The new caps are observable behaviour: under sustained load the Map/Set sizes plateau at the configured maxima rather than monotone-growing.
|
||||
新上限是可观察行为:在持续负载下,Map/Set 大小在配置的最大值处趋于平稳,而不是单调增长。
|
||||
|
||||
## 关联笔记
|
||||
|
||||
- [[sur-loop-scheduled-oom]]
|
||||
|
||||
Reference in New Issue
Block a user