{ "topic": "heavykeeper-topk", "type": "code_reading", "schema_version": "1.0.0", "generated": "2026-09-09T21:42:00+08:00", "questions": [ { "id": "cr-001", "type": "code_reading", "difficulty": 3, "tags": [ "heavykeeper-topk", "heavykeeper" ], "question": "阅读以下 HeavyKeeper.add() 方法的核心逻辑,回答子问题。", "code": "public AddResult add(String key, int increment) {\n byte[] keyBytes = key.getBytes();\n long itemFingerprint = hash(keyBytes);\n int maxCount = 0;\n for (int i = 0; i < depth; i++) {\n int bucketNumber = Math.abs(hash(keyBytes)) % width;\n Bucket bucket = buckets[i][bucketNumber];\n synchronized (bucket) {\n if (bucket.count == 0) {\n bucket.fingerprint = itemFingerprint;\n bucket.count = increment;\n } else if (bucket.fingerprint == itemFingerprint) {\n bucket.count += increment;\n } else {\n for (int j = 0; j < increment; j++) {\n double decay = bucket.count < LOOKUP_TABLE_SIZE ?\n lookupTable[bucket.count] :\n lookupTable[LOOKUP_TABLE_SIZE - 1];\n if (random.nextDouble() < decay) {\n bucket.count--;\n if (bucket.count == 0) {\n bucket.fingerprint = itemFingerprint;\n bucket.count = increment - j;\n break;\n }\n }\n }\n }\n }\n maxCount = Math.max(maxCount, bucket.count);\n }\n total += increment;\n return new AddResult(maxCount, ...);\n}", "language": "java", "explanation": "这段代码展示了 HeavyKeeper 的核心插入逻辑。对于每个 Key,在 depth 层中分别定位桶,然后根据桶的状态(空/匹配/冲突)执行不同操作。冲突时通过概率衰减尝试腾空桶。", "source": null, "related": [], "sub_questions": [ { "index": 1, "type": "single_choice", "question": "当 `bucket.fingerprint == itemFingerprint` 为真时,说明什么?", "options": { "A": "发生了哈希冲突,需要进行衰减淘汰", "B": "该桶当前存储的正是当前 Key,直接累加计数", "C": "桶已被清空,可以写入新指纹", "D": "需要将该 Key 移到下一层的桶中" }, "answer": "B", "explanation": "bucket.fingerprint == itemFingerprint 表示桶中已存储的指纹与当前 Key 的指纹一致,说明该桶属于当前 Key,此时直接将计数累加 increment,无需任何衰减操作。这是最理想的分支——热 Key 被再次命中。" }, { "index": 2, "type": "single_choice", "question": "在冲突衰减分支中,decay 值的计算方式为 `lookupTable[bucket.count]`(当 count < 256 时),这意味着桶计数越高,衰减概率如何变化?", "options": { "A": "衰减概率越大,更容易被淘汰", "B": "衰减概率越小,越难被淘汰", "C": "衰减概率保持不变", "D": "衰减概率先增后减" }, "answer": "B", "explanation": "lookupTable[i] = 0.92^i,这是一个单调递减函数。bucket.count 越大,lookupTable[bucket.count] 越小,意味着递减概率越低。因此高频热 Key 的桶计数大、衰减概率小,很难被冷 Key 冲掉;而冷 Key 的桶计数小、衰减概率大,容易被清零淘汰。" }, { "index": 3, "type": "single_choice", "question": "以下哪个场景会触发 `bucket.count = increment - j` 这行代码?", "options": { "A": "桶为空时直接写入新 Key", "B": "当前 Key 的指纹与桶指纹匹配时累加计数", "C": "冲突衰减过程中桶计数被减到 0,桶被腾空", "D": "minHeap 中的最小元素被淘汰时" }, "answer": "C", "explanation": "bucket.count == 0 发生在冲突衰减循环中:经过 j 次成功的衰减递减后,桶计数从原来的状态被减到了 0。此时桶已腾空,写入当前 Key 的指纹,并将 count 设为 increment - j(因为已经衰减了 j 次,还剩 increment - j 次未使用)。break 跳出衰减循环。" } ] }, { "id": "cr-002", "type": "code_reading", "difficulty": 4, "tags": [ "heavykeeper-topk", "heavykeeper" ], "question": "阅读以下 HeavyKeeper 的 add() 方法完整片段(含 minHeap 管理),回答子问题。", "code": "// ... 前面的桶操作逻辑 ...\n// 经过 depth 层桶操作后,得到 maxCount\n\ntotal += increment;\n\n// === minHeap 管理 ===\nsynchronized (minHeap) {\n if (maxCount >= minCount) {\n // 检查 Key 是否已在堆中\n Node existing = heapIndex.get(key);\n if (existing != null) {\n existing.count = maxCount;\n minHeap.remove(existing);\n minHeap.add(existing);\n } else {\n if (minHeap.size() < k) {\n Node node = new Node(key, maxCount);\n minHeap.add(node);\n heapIndex.put(key, node);\n } else {\n Node min = minHeap.peek();\n if (maxCount > min.count) {\n heapIndex.remove(min.key);\n minHeap.poll();\n Node node = new Node(key, maxCount);\n minHeap.add(node);\n heapIndex.put(key, node);\n }\n }\n }\n }\n}", "language": "java", "explanation": "这段代码展示了 minHeap 的完整管理逻辑:先检查 maxCount 是否超过 minCount 门槛,再判断 Key 是否已在堆中,最后根据堆大小决定是直接插入还是替换堆顶最小元素。", "source": null, "related": [], "sub_questions": [ { "index": 1, "type": "single_choice", "question": "为什么要检查 `maxCount >= minCount` 才能进入 minHeap 操作?", "options": { "A": "防止整数溢出", "B": "过滤低频 Key,避免大量低频元素占用堆空间", "C": "确保 fingerprint 已初始化", "D": "保证哈希分布均匀" }, "answer": "B", "explanation": "minCount 是一个入门门槛。数据流中绝大多数 Key 是低频的,如果每个 Key 都加入堆,堆的大小会无限膨胀。只有频率达到 minCount 的 Key 才值得参与 Top-K 竞争,这大大减少了堆的操作次数和内存占用。" }, { "index": 2, "type": "single_choice", "question": "当 minHeap.size() == k 且 maxCount > min.count 时,代码执行了什么操作?", "options": { "A": "在堆中插入新节点,堆大小变为 k+1", "B": "移除堆顶(最小元素),将新 Key 插入堆中,堆大小保持为 k", "C": "忽略新 Key,不做任何操作", "D": "将堆顶元素的 count 更新为 maxCount" }, "answer": "B", "explanation": "当堆已满(size == k)且新 Key 的频率大于堆顶最小元素时,先 poll() 移除堆顶,再 add() 插入新 Key。这样堆大小保持为 k,同时将更高频的 Key 纳入 Top-K。这是维护最小堆实现 Top-K 的标准操作。" }, { "index": 3, "type": "single_choice", "question": "heapIndex(Map)的作用是什么?", "options": { "A": "存储所有 Key 的指纹", "B": "记录 Key 到堆节点的映射,支持快速查找和更新已在堆中的 Key", "C": "记录每层桶的访问频率", "D": "缓存 fading() 操作的中间结果" }, "answer": "B", "explanation": "Java 的 PriorityQueue 不支持高效的随机查找和删除。heapIndex 维护了 Key → Node 的映射,当一个已在堆中的 Key 频率更新时,可以通过 heapIndex 快速找到对应 Node,执行 remove + add 更新堆。如果不维护这个索引,每次更新都需要 O(n) 遍历堆。" } ] }, { "id": "cr-003", "type": "code_reading", "difficulty": 5, "tags": [ "heavykeeper-topk", "heavykeeper" ], "question": "阅读以下 HeavyKeeper.fading() 方法的完整实现,回答子问题。", "code": "public void fading() {\n // 第一阶段:对所有桶执行右移减半\n for (Bucket[] row : buckets) {\n for (Bucket bucket : row) {\n synchronized (bucket) {\n bucket.count = bucket.count >> 1;\n }\n }\n }\n // 第二阶段:同步衰减 minHeap 中所有节点\n synchronized (minHeap) {\n PriorityQueue newHeap = new PriorityQueue<>(\n k, Comparator.comparingLong(n -> n.count)\n );\n for (Node node : minHeap) {\n newHeap.add(new Node(node.key, node.count >> 1));\n }\n minHeap.clear();\n minHeap.addAll(newHeap);\n }\n // 第三阶段:全局计数减半\n total = total >> 1;\n}", "language": "java", "explanation": "fading() 实现了全局指数衰减,通过三阶段操作:桶计数减半、堆节点计数减半、全局 total 减半。这使 HeavyKeeper 能适应数据流分布的时间变化。", "source": null, "related": [], "sub_questions": [ { "index": 1, "type": "single_choice", "question": "fading() 为什么要新建一个 newHeap 而不是直接修改 minHeap 中的 Node.count?", "options": { "A": "因为 Node.count 是 final 字段,不可修改", "B": "因为直接修改 Node 的 count 会破坏 PriorityQueue 的堆序性质,导致堆结构失效", "C": "为了提高内存分配效率", "D": "因为 minHeap 不支持迭代器" }, "answer": "B", "explanation": "PriorityQueue 是基于堆序(每个父节点 ≤ 子节点)维护的。如果直接修改 Node 的 count 值,堆中元素的顺序关系会被破坏,后续的 peek/poll 操作会返回错误结果。正确做法是创建新堆,将衰减后的节点逐个插入,重新建立堆序。" }, { "index": 2, "type": "single_choice", "question": "如果某个 Key 的桶 count 在 fading() 前为 3,执行 fading() 后变为多少?", "options": { "A": "3", "B": "1", "C": "0", "D": "1.5" }, "answer": "B", "explanation": "3 >> 1 = 1(二进制 11 右移一位变成 1,即整数除以 2 向下取整)。fading() 后该桶的 count 从 3 变为 1。如果原来是 1,则 1 >> 1 = 0,桶被清空。这种衰减让低频 Key 逐渐被淘汰。" }, { "index": 3, "type": "single_choice", "question": "fading() 执行后,一个频率为 100 的热 Key 经过多次 fading() 后频率变为约 12(≈100/8),这意味着大约执行了几次 fading()?", "options": { "A": "2 次", "B": "3 次", "C": "4 次", "D": "5 次" }, "answer": "B", "explanation": "每次 fading() 将计数减半:100 → 50 → 25 → 12,经过 3 次 fading() 后频率约为 12(精确值为 12,因为 100>>1=50, 50>>1=25, 25>>1=12)。这体现了 fading 的指数衰减效果——频率每经过一次 fading() 就减半。" } ] } ] }