Files
examination/topics/thumbup/heavykeeper-topk/code_reading.json
T

176 lines
12 KiB
JSON
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
{
"topic": "heavykeeper-topk",
"type": "code_reading",
"schema_version": "1.0.0",
"generated": "2026-09-09T21:42:00+08:00",
"questions": [
{
"id": "cr-001",
"type": "code_reading",
"difficulty": 3,
"tags": [
"heavykeeper-topk",
"heavykeeper"
],
"question": "阅读以下 HeavyKeeper.add() 方法的核心逻辑,回答子问题。",
"code": "public AddResult add(String key, int increment) {\n byte[] keyBytes = key.getBytes();\n long itemFingerprint = hash(keyBytes);\n int maxCount = 0;\n for (int i = 0; i < depth; i++) {\n int bucketNumber = Math.abs(hash(keyBytes)) % width;\n Bucket bucket = buckets[i][bucketNumber];\n synchronized (bucket) {\n if (bucket.count == 0) {\n bucket.fingerprint = itemFingerprint;\n bucket.count = increment;\n } else if (bucket.fingerprint == itemFingerprint) {\n bucket.count += increment;\n } else {\n for (int j = 0; j < increment; j++) {\n double decay = bucket.count < LOOKUP_TABLE_SIZE ?\n lookupTable[bucket.count] :\n lookupTable[LOOKUP_TABLE_SIZE - 1];\n if (random.nextDouble() < decay) {\n bucket.count--;\n if (bucket.count == 0) {\n bucket.fingerprint = itemFingerprint;\n bucket.count = increment - j;\n break;\n }\n }\n }\n }\n }\n maxCount = Math.max(maxCount, bucket.count);\n }\n total += increment;\n return new AddResult(maxCount, ...);\n}",
"language": "java",
"explanation": "这段代码展示了 HeavyKeeper 的核心插入逻辑。对于每个 Key,在 depth 层中分别定位桶,然后根据桶的状态(空/匹配/冲突)执行不同操作。冲突时通过概率衰减尝试腾空桶。",
"source": null,
"related": [],
"sub_questions": [
{
"index": 1,
"type": "single_choice",
"question": "当 `bucket.fingerprint == itemFingerprint` 为真时,说明什么?",
"options": {
"A": "发生了哈希冲突,需要进行衰减淘汰",
"B": "该桶当前存储的正是当前 Key,直接累加计数",
"C": "桶已被清空,可以写入新指纹",
"D": "需要将该 Key 移到下一层的桶中"
},
"answer": "B",
"explanation": "bucket.fingerprint == itemFingerprint 表示桶中已存储的指纹与当前 Key 的指纹一致,说明该桶属于当前 Key,此时直接将计数累加 increment,无需任何衰减操作。这是最理想的分支——热 Key 被再次命中。"
},
{
"index": 2,
"type": "single_choice",
"question": "在冲突衰减分支中,decay 值的计算方式为 `lookupTable[bucket.count]`(当 count < 256 时),这意味着桶计数越高,衰减概率如何变化?",
"options": {
"A": "衰减概率越大,更容易被淘汰",
"B": "衰减概率越小,越难被淘汰",
"C": "衰减概率保持不变",
"D": "衰减概率先增后减"
},
"answer": "B",
"explanation": "lookupTable[i] = 0.92^i,这是一个单调递减函数。bucket.count 越大,lookupTable[bucket.count] 越小,意味着递减概率越低。因此高频热 Key 的桶计数大、衰减概率小,很难被冷 Key 冲掉;而冷 Key 的桶计数小、衰减概率大,容易被清零淘汰。"
},
{
"index": 3,
"type": "single_choice",
"question": "以下哪个场景会触发 `bucket.count = increment - j` 这行代码?",
"options": {
"A": "桶为空时直接写入新 Key",
"B": "当前 Key 的指纹与桶指纹匹配时累加计数",
"C": "冲突衰减过程中桶计数被减到 0,桶被腾空",
"D": "minHeap 中的最小元素被淘汰时"
},
"answer": "C",
"explanation": "bucket.count == 0 发生在冲突衰减循环中:经过 j 次成功的衰减递减后,桶计数从原来的状态被减到了 0。此时桶已腾空,写入当前 Key 的指纹,并将 count 设为 increment - j(因为已经衰减了 j 次,还剩 increment - j 次未使用)。break 跳出衰减循环。"
}
]
},
{
"id": "cr-002",
"type": "code_reading",
"difficulty": 4,
"tags": [
"heavykeeper-topk",
"heavykeeper"
],
"question": "阅读以下 HeavyKeeper 的 add() 方法完整片段(含 minHeap 管理),回答子问题。",
"code": "// ... 前面的桶操作逻辑 ...\n// 经过 depth 层桶操作后,得到 maxCount\n\ntotal += increment;\n\n// === minHeap 管理 ===\nsynchronized (minHeap) {\n if (maxCount >= minCount) {\n // 检查 Key 是否已在堆中\n Node existing = heapIndex.get(key);\n if (existing != null) {\n existing.count = maxCount;\n minHeap.remove(existing);\n minHeap.add(existing);\n } else {\n if (minHeap.size() < k) {\n Node node = new Node(key, maxCount);\n minHeap.add(node);\n heapIndex.put(key, node);\n } else {\n Node min = minHeap.peek();\n if (maxCount > min.count) {\n heapIndex.remove(min.key);\n minHeap.poll();\n Node node = new Node(key, maxCount);\n minHeap.add(node);\n heapIndex.put(key, node);\n }\n }\n }\n }\n}",
"language": "java",
"explanation": "这段代码展示了 minHeap 的完整管理逻辑:先检查 maxCount 是否超过 minCount 门槛,再判断 Key 是否已在堆中,最后根据堆大小决定是直接插入还是替换堆顶最小元素。",
"source": null,
"related": [],
"sub_questions": [
{
"index": 1,
"type": "single_choice",
"question": "为什么要检查 `maxCount >= minCount` 才能进入 minHeap 操作?",
"options": {
"A": "防止整数溢出",
"B": "过滤低频 Key,避免大量低频元素占用堆空间",
"C": "确保 fingerprint 已初始化",
"D": "保证哈希分布均匀"
},
"answer": "B",
"explanation": "minCount 是一个入门门槛。数据流中绝大多数 Key 是低频的,如果每个 Key 都加入堆,堆的大小会无限膨胀。只有频率达到 minCount 的 Key 才值得参与 Top-K 竞争,这大大减少了堆的操作次数和内存占用。"
},
{
"index": 2,
"type": "single_choice",
"question": "当 minHeap.size() == k 且 maxCount > min.count 时,代码执行了什么操作?",
"options": {
"A": "在堆中插入新节点,堆大小变为 k+1",
"B": "移除堆顶(最小元素),将新 Key 插入堆中,堆大小保持为 k",
"C": "忽略新 Key,不做任何操作",
"D": "将堆顶元素的 count 更新为 maxCount"
},
"answer": "B",
"explanation": "当堆已满(size == k)且新 Key 的频率大于堆顶最小元素时,先 poll() 移除堆顶,再 add() 插入新 Key。这样堆大小保持为 k,同时将更高频的 Key 纳入 Top-K。这是维护最小堆实现 Top-K 的标准操作。"
},
{
"index": 3,
"type": "single_choice",
"question": "heapIndex(Map<String, Node>)的作用是什么?",
"options": {
"A": "存储所有 Key 的指纹",
"B": "记录 Key 到堆节点的映射,支持快速查找和更新已在堆中的 Key",
"C": "记录每层桶的访问频率",
"D": "缓存 fading() 操作的中间结果"
},
"answer": "B",
"explanation": "Java 的 PriorityQueue 不支持高效的随机查找和删除。heapIndex 维护了 Key → Node 的映射,当一个已在堆中的 Key 频率更新时,可以通过 heapIndex 快速找到对应 Node,执行 remove + add 更新堆。如果不维护这个索引,每次更新都需要 O(n) 遍历堆。"
}
]
},
{
"id": "cr-003",
"type": "code_reading",
"difficulty": 5,
"tags": [
"heavykeeper-topk",
"heavykeeper"
],
"question": "阅读以下 HeavyKeeper.fading() 方法的完整实现,回答子问题。",
"code": "public void fading() {\n // 第一阶段:对所有桶执行右移减半\n for (Bucket[] row : buckets) {\n for (Bucket bucket : row) {\n synchronized (bucket) {\n bucket.count = bucket.count >> 1;\n }\n }\n }\n // 第二阶段:同步衰减 minHeap 中所有节点\n synchronized (minHeap) {\n PriorityQueue<Node> newHeap = new PriorityQueue<>(\n k, Comparator.comparingLong(n -> n.count)\n );\n for (Node node : minHeap) {\n newHeap.add(new Node(node.key, node.count >> 1));\n }\n minHeap.clear();\n minHeap.addAll(newHeap);\n }\n // 第三阶段:全局计数减半\n total = total >> 1;\n}",
"language": "java",
"explanation": "fading() 实现了全局指数衰减,通过三阶段操作:桶计数减半、堆节点计数减半、全局 total 减半。这使 HeavyKeeper 能适应数据流分布的时间变化。",
"source": null,
"related": [],
"sub_questions": [
{
"index": 1,
"type": "single_choice",
"question": "fading() 为什么要新建一个 newHeap 而不是直接修改 minHeap 中的 Node.count?",
"options": {
"A": "因为 Node.count 是 final 字段,不可修改",
"B": "因为直接修改 Node 的 count 会破坏 PriorityQueue 的堆序性质,导致堆结构失效",
"C": "为了提高内存分配效率",
"D": "因为 minHeap 不支持迭代器"
},
"answer": "B",
"explanation": "PriorityQueue 是基于堆序(每个父节点 ≤ 子节点)维护的。如果直接修改 Node 的 count 值,堆中元素的顺序关系会被破坏,后续的 peek/poll 操作会返回错误结果。正确做法是创建新堆,将衰减后的节点逐个插入,重新建立堆序。"
},
{
"index": 2,
"type": "single_choice",
"question": "如果某个 Key 的桶 count 在 fading() 前为 3,执行 fading() 后变为多少?",
"options": {
"A": "3",
"B": "1",
"C": "0",
"D": "1.5"
},
"answer": "B",
"explanation": "3 >> 1 = 1(二进制 11 右移一位变成 1,即整数除以 2 向下取整)。fading() 后该桶的 count 从 3 变为 1。如果原来是 1,则 1 >> 1 = 0,桶被清空。这种衰减让低频 Key 逐渐被淘汰。"
},
{
"index": 3,
"type": "single_choice",
"question": "fading() 执行后,一个频率为 100 的热 Key 经过多次 fading() 后频率变为约 12(≈100/8),这意味着大约执行了几次 fading()?",
"options": {
"A": "2 次",
"B": "3 次",
"C": "4 次",
"D": "5 次"
},
"answer": "B",
"explanation": "每次 fading() 将计数减半:100 → 50 → 25 → 12,经过 3 次 fading() 后频率约为 12(精确值为 12,因为 100>>1=50, 50>>1=25, 25>>1=12)。这体现了 fading 的指数衰减效果——频率每经过一次 fading() 就减半。"
}
]
}
]
}