0f68a64829
Deploy Examination / deploy (push) Successful in 10s
Subtopics: - distributed-microservice: 45 questions (分布式微服务架构) - message-queue: 45 questions (消息队列) - k8s-observability: 45 questions (K8s与可观测性) - go-java-concurrency: 45 questions (Go/Java并发模型) - database-advanced: 35 questions (数据库进阶) - ai-engineering: 35 questions (AI工程实践) Question types: single_choice, true_false, fill_blank, short_answer, code_reading
309 lines
27 KiB
JSON
309 lines
27 KiB
JSON
{
|
||
"topic": "k8s-observability",
|
||
"type": "code_reading",
|
||
"schema_version": "1.0.0",
|
||
"generated": "2026-09-09T16:10:00+08:00",
|
||
"questions": [
|
||
{
|
||
"id": "cr-001",
|
||
"type": "code_reading",
|
||
"difficulty": 3,
|
||
"tags": [
|
||
"kubernetes",
|
||
"deployment",
|
||
"rolling-update",
|
||
"liveness",
|
||
"readiness"
|
||
],
|
||
"question": "以下是一个微服务的 Deployment YAML 配置,阅读后回答问题。",
|
||
"code": "apiVersion: apps/v1\nkind: Deployment\nmetadata:\n name: order-service\n namespace: production\n labels:\n app: order-service\n version: v2\nspec:\n replicas: 4\n strategy:\n type: RollingUpdate\n rollingUpdate:\n maxSurge: 1\n maxUnavailable: 0\n selector:\n matchLabels:\n app: order-service\n template:\n metadata:\n labels:\n app: order-service\n version: v2\n spec:\n containers:\n - name: order\n image: registry.example.com/order-service:2.3.1\n ports:\n - containerPort: 8080\n resources:\n requests:\n cpu: 200m\n memory: 256Mi\n limits:\n cpu: 1000m\n memory: 512Mi\n livenessProbe:\n httpGet:\n path: /health/live\n port: 8080\n initialDelaySeconds: 15\n periodSeconds: 10\n failureThreshold: 3\n readinessProbe:\n httpGet:\n path: /health/ready\n port: 8080\n initialDelaySeconds: 5\n periodSeconds: 5\n failureThreshold: 2\n env:\n - name: DB_HOST\n valueFrom:\n configMapKeyRef:\n name: order-config\n key: db_host\n terminationGracePeriodSeconds: 30",
|
||
"language": "yaml",
|
||
"explanation": "本题考察 K8s Deployment 的核心配置,包括滚动更新策略、资源限制、健康检查探针等关键概念。需要理解 maxSurge/maxUnavailable 的语义以及 liveness 与 readiness 探针的区别。",
|
||
"source": null,
|
||
"related": [],
|
||
"sub_questions": [
|
||
{
|
||
"index": 1,
|
||
"type": "single_choice",
|
||
"question": "当前 Deployment 配置了 maxSurge=1、maxUnavailable=0 的滚动更新策略。在执行滚动更新时,Kubernetes 会先做什么?",
|
||
"options": {
|
||
"A": "先终止一个旧 Pod,再创建一个新 Pod",
|
||
"B": "先创建一个新 Pod(达到 replicas+1=5),待其就绪后再逐步替换旧 Pod",
|
||
"C": "同时终止所有旧 Pod 并创建所有新 Pod",
|
||
"D": "创建两个新 Pod,然后再终止两个旧 Pod"
|
||
},
|
||
"answer": "B",
|
||
"explanation": "maxSurge=1 表示滚动更新过程中最多可以比期望副本数多出 1 个 Pod(即总共 5 个 Pod)。maxUnavailable=0 表示更新期间不允许任何 Pod 不可用。因此 K8s 会先创建一个新 Pod,等其通过 readinessProbe 变为 Ready 后,才终止一个旧 Pod,如此循环直到所有 Pod 更新完毕。这种策略保证了零停机滚动更新。"
|
||
},
|
||
{
|
||
"index": 2,
|
||
"type": "single_choice",
|
||
"question": "如果 order-service 的新版本 Pod 因启动时间较长,livenessProbe 在 initialDelaySeconds=15 后开始探测,探测 3 次失败后 Pod 会被重启。此时 readinessProbe 的 initialDelaySeconds=5 和 periodSeconds=5 意味着什么?",
|
||
"options": {
|
||
"A": "readinessProbe 不会影响 Pod 的重启",
|
||
"B": "readinessProbe 在 Pod 启动 5 秒后开始,每 5 秒探测一次,连续 2 次失败则 Pod 被标记为 NotReady 并从 Service 端点移除",
|
||
"C": "readinessProbe 只在首次探测失败后才每 5 秒探测一次",
|
||
"D": "readinessProbe 和 livenessProbe 的失败次数阈值相同时行为一致"
|
||
},
|
||
"answer": "B",
|
||
"explanation": "readinessProbe 的 initialDelaySeconds=5 表示 Pod 启动后 5 秒开始探测,periodSeconds=5 表示每 5 秒探测一次,failureThreshold=2 表示连续 2 次失败后 Pod 会被标记为 NotReady。NotReady 的 Pod 不会接收来自 Service 的流量(从 Endpoint 列表中移除),但不会被重启。这与 livenessProbe(失败后重启 Pod)的语义完全不同。readinessProbe 确保只有真正准备好接收请求的 Pod 才会被分发流量。"
|
||
},
|
||
{
|
||
"index": 3,
|
||
"type": "short_answer",
|
||
"question": "该 Deployment 的 resources 配置中,requests 和 limits 的 CPU 值分别是多少?如果某个节点上已有其他 Pod 占用了大部分 CPU,新的 order-service Pod 可能出现什么问题?",
|
||
"answer": "requests.cpu=200m, limits.cpu=1000m。如果节点 CPU 资源不足,新 Pod 可能处于 Pending 状态无法调度;即使调度成功,在 CPU 使用量接近 1000m 时会被内核限流(throttle),导致延迟升高。",
|
||
"keywords": [
|
||
"200m",
|
||
"1000m",
|
||
"Pending",
|
||
"throttle",
|
||
"CPU 限流"
|
||
],
|
||
"scoring_rubric": "答出 requests.cpu=200m(1分)、limits.cpu=1000m(1分)、节点资源不足导致 Pending(1分)、CPU throttle 现象(1分),共4分。",
|
||
"explanation": "requests.cpu=200m 表示 Pod 启动时至少需要 0.2 个 CPU 核心的保证资源,limits.cpu=1000m 表示最多可使用 1 个 CPU 核心。当节点上可分配的 CPU 资源少于 200m 时,新 Pod 无法被调度(Pending)。即使调度成功,当 Pod 的 CPU 使用接近 limits 时,Linux CFS 调度器会对进程进行 throttle(限流),导致请求处理延迟增加。requests 用于调度决策,limits 用于运行时资源上限控制。"
|
||
}
|
||
]
|
||
},
|
||
{
|
||
"id": "cr-002",
|
||
"type": "code_reading",
|
||
"difficulty": 3,
|
||
"tags": [
|
||
"kubernetes",
|
||
"HPA",
|
||
"pod",
|
||
"deployment"
|
||
],
|
||
"question": "以下是一个 Horizontal Pod Autoscaler (HPA) 的配置,阅读后回答问题。",
|
||
"code": "apiVersion: autoscaling/v2\nkind: HorizontalPodAutoscaler\nmetadata:\n name: order-service-hpa\n namespace: production\nspec:\n scaleTargetRef:\n apiVersion: apps/v1\n kind: Deployment\n name: order-service\n minReplicas: 3\n maxReplicas: 20\n metrics:\n - type: Resource\n resource:\n name: cpu\n target:\n type: Utilization\n averageUtilization: 65\n - type: Resource\n resource:\n name: memory\n target:\n type: Utilization\n averageUtilization: 80\n behavior:\n scaleUp:\n stabilizationWindowSeconds: 60\n policies:\n - type: Pods\n value: 4\n periodSeconds: 60\n scaleDown:\n stabilizationWindowSeconds: 300\n policies:\n - type: Percent\n value: 10\n periodSeconds: 60",
|
||
"language": "yaml",
|
||
"explanation": "本题考察 HPA 的配置细节,包括多指标扩缩容、behavior 策略(稳定窗口、扩缩速度控制)等高级特性。需要理解 HPA 如何综合多个指标进行决策。",
|
||
"source": null,
|
||
"related": [],
|
||
"sub_questions": [
|
||
{
|
||
"index": 1,
|
||
"type": "single_choice",
|
||
"question": "该 HPA 配置了 CPU 平均利用率 65% 和内存平均利用率 80% 两个指标。当 CPU 利用率升至 75% 但内存利用率仅 40% 时,HPA 会如何决策?",
|
||
"options": {
|
||
"A": "不扩容,因为内存利用率未达到 80% 的阈值",
|
||
"B": "扩容,HPA 取各指标所需副本数的最大值",
|
||
"C": "扩容,HPA 取各指标所需副本数的平均值",
|
||
"D": "扩容,HPA 取各指标所需副本数的最小值"
|
||
},
|
||
"answer": "B",
|
||
"explanation": "HPA 使用多个指标时,会分别计算每个指标所期望的副本数,然后取所有指标中需要副本数最多的那个作为最终目标副本数(取最大值)。CPU 利用率 75% 超过 65% 的目标值,HPA 会计算出一个更大的副本数;即使内存利用率低于目标,HPA 仍会以 CPU 指标计算出的较大值为准进行扩容。这种保守策略确保所有指标都在目标范围内。"
|
||
},
|
||
{
|
||
"index": 2,
|
||
"type": "single_choice",
|
||
"question": "关于 behavior 配置,scaleDown 的 stabilizationWindowSeconds=300 和 policies 中 value=10 的组合意味着什么?",
|
||
"options": {
|
||
"A": "缩容时立即缩减 10 个 Pod",
|
||
"B": "在 300 秒稳定窗口内,如果指标持续低于目标,每 60 秒最多缩减当前副本数的 10%",
|
||
"C": "缩容速度比扩容快 5 倍",
|
||
"D": "稳定窗口期间不允许任何缩容操作"
|
||
},
|
||
"answer": "B",
|
||
"explanation": "scaleDown 的 stabilizationWindowSeconds=300 表示 HPA 在决定缩容时,会回看过去 300 秒内的最小推荐副本数,只有当当前目标副本数低于这个最小值时才会执行缩容。policies 中 type=Percent、value=10、periodSeconds=60 表示在 60 秒的时间窗口内,最多缩减当前副本数的 10%。例如当前有 10 个 Pod,每分钟最多缩 1 个。这种机制避免了指标短暂波动导致的频繁缩容。"
|
||
},
|
||
{
|
||
"index": 3,
|
||
"type": "short_answer",
|
||
"question": "该 HPA 的 scaleUp 和 scaleDown 配置在稳定窗口和策略上有什么不对称设计?这种设计的目的是什么?",
|
||
"answer": "scaleUp 稳定窗口 60 秒、每分钟最多增 4 个 Pod(较快);scaleDown 稳定窗口 300 秒、每分钟最多缩 10%(较慢)。目的是快速响应流量高峰、缓慢收缩以避免缩容过快导致的服务抖动。",
|
||
"keywords": [
|
||
"60秒",
|
||
"300秒",
|
||
"快速扩容",
|
||
"缓慢缩容",
|
||
"抖动",
|
||
"稳定性"
|
||
],
|
||
"scoring_rubric": "答出 scaleUp 更激进(1分)、scaleDown 更保守(1分)、具体数值对比(1分)、解释快扩慢缩的工程目的(1分),共4分。",
|
||
"explanation": "这是一个典型的「快扩慢缩」设计模式。scaleUp 窗口短(60s)、允许单次增加较多 Pod(最多 4 个),以便快速应对流量突增;scaleDown 窗口长(300s)、单次缩减比例小(10%),防止因指标短暂波动或流量临时下降而过度缩容导致服务不可用。在生产环境中,缩容通常比扩容更危险——缩太快可能导致请求超载或丢失连接,因此 HPA 默认和最佳实践都推荐对缩容采取更保守的策略。"
|
||
}
|
||
]
|
||
},
|
||
{
|
||
"id": "cr-003",
|
||
"type": "code_reading",
|
||
"difficulty": 2,
|
||
"tags": [
|
||
"kubernetes",
|
||
"service",
|
||
"ingress"
|
||
],
|
||
"question": "以下是微服务的 Service 和 Ingress 配置,阅读后回答问题。",
|
||
"code": "apiVersion: v1\nkind: Service\nmetadata:\n name: order-service\n namespace: production\n labels:\n app: order-service\nspec:\n type: ClusterIP\n ports:\n - port: 80\n targetPort: 8080\n protocol: TCP\n selector:\n app: order-service\n---\napiVersion: networking.k8s.io/v1\nkind: Ingress\nmetadata:\n name: api-gateway\n namespace: production\n annotations:\n nginx.ingress.kubernetes.io/ssl-redirect: \"true\"\n nginx.ingress.kubernetes.io/proxy-body-size: \"10m\"\nspec:\n ingressClassName: nginx\n tls:\n - hosts:\n - api.example.com\n secretName: api-tls-secret\n rules:\n - host: api.example.com\n http:\n paths:\n - path: /api/v1/orders\n pathType: Prefix\n backend:\n service:\n name: order-service\n port:\n number: 80\n - path: /api/v1/products\n pathType: Prefix\n backend:\n service:\n name: product-service\n port:\n number: 80",
|
||
"language": "yaml",
|
||
"explanation": "本题考察 K8s Service 与 Ingress 的配置和路由规则。需要理解 ClusterIP Service 的作用、Ingress 的路径路由机制以及 TLS 终止等概念。",
|
||
"source": null,
|
||
"related": [],
|
||
"sub_questions": [
|
||
{
|
||
"index": 1,
|
||
"type": "single_choice",
|
||
"question": "Service 的 targetPort=8080 而 port=80,这意味着什么?如果客户端通过 kubectl exec 进入集群内任意 Pod 并 curl order-service:80,请求会被转发到哪里?",
|
||
"options": {
|
||
"A": "请求到达 order-service Pod 的 80 端口",
|
||
"B": "请求到达 order-service Pod 的 8080 端口",
|
||
"C": "请求会被拒绝,因为集群内不能用 Service 名称访问",
|
||
"D": "请求到达 Ingress controller 的 80 端口"
|
||
},
|
||
"answer": "B",
|
||
"explanation": "port=80 是 Service 暴露给集群内部的端口,targetPort=8080 是后端 Pod 实际监听的端口。当集群内的客户端访问 order-service:80 时,kube-proxy 会将流量转发到 Pod 的 8080 端口。这就是 Service 的端口映射机制——通过统一的 Service 端口访问不同后端容器端口的应用。"
|
||
},
|
||
{
|
||
"index": 2,
|
||
"type": "single_choice",
|
||
"question": "用户通过浏览器访问 https://api.example.com/api/v1/orders/123,该请求的完整转发链路是什么?",
|
||
"options": {
|
||
"A": "浏览器 → Ingress (443) → order-service ClusterIP (80) → Pod (8080)",
|
||
"B": "浏览器 → Ingress (80) → order-service ClusterIP (80) → Pod (8080)",
|
||
"C": "浏览器 → Ingress (443) → Pod (8080),跳过了 Service 层",
|
||
"D": "浏览器 → Ingress (443) → product-service ClusterIP (80) → Pod (8080)"
|
||
},
|
||
"answer": "A",
|
||
"explanation": "请求链路为:浏览器发起 HTTPS 请求 → Ingress Controller 监听 443 端口并进行 TLS 终止(解密后变为 HTTP)→ 根据 path=/api/v1/orders 和 host=api.example.com 匹配路由规则 → 转发到 order-service ClusterIP Service 的 80 端口 → kube-proxy 将流量转发到 Pod 的 8080 端口。注意 TLS 终止发生在 Ingress 层,Ingress 到 Service 之间是明文 HTTP。"
|
||
},
|
||
{
|
||
"index": 3,
|
||
"type": "single_choice",
|
||
"question": "如果此时访问 https://api.example.com/api/v1/orders 会得到什么结果?(假设 order-service Pod 正常运行)",
|
||
"options": {
|
||
"A": "返回 404,因为路径 /api/v1/orders 不在配置中",
|
||
"B": "成功到达 order-service,因为 pathType: Prefix 会匹配以 /api/v1/orders 开头的所有路径",
|
||
"C": "返回 403 Forbidden",
|
||
"D": "请求被路由到 product-service"
|
||
},
|
||
"answer": "B",
|
||
"explanation": "pathType: Prefix 表示前缀匹配,/api/v1/orders 会匹配所有以该路径开头的请求,包括 /api/v1/orders、/api/v1/orders/123、/api/v1/orders/search 等。如果使用 pathType: Exact,则只有完全等于 /api/v1/orders 的请求才会被匹配。在生产环境中,需要根据 API 设计选择合适的 pathType 以避免路由冲突。"
|
||
}
|
||
]
|
||
},
|
||
{
|
||
"id": "cr-004",
|
||
"type": "code_reading",
|
||
"difficulty": 3,
|
||
"tags": [
|
||
"prometheus",
|
||
"grafana",
|
||
"kubernetes",
|
||
"service"
|
||
],
|
||
"question": "以下是一组 Prometheus 监控配置,包括 ServiceMonitor 和告警规则,阅读后回答问题。",
|
||
"code": "apiVersion: monitoring.coreos.com/v1\nkind: ServiceMonitor\nmetadata:\n name: order-service-monitor\n namespace: production\n labels:\n release: prometheus\nspec:\n selector:\n matchLabels:\n app: order-service\n endpoints:\n - port: metrics\n interval: 30s\n path: /metrics\n namespaceSelector:\n matchNames:\n - production\n---\ngroups:\n- name: order-service-alerts\n rules:\n - alert: HighErrorRate\n expr: |\n sum(rate(http_requests_total{app=\"order-service\", status=~\"5..\"}[5m]))\n /\n sum(rate(http_requests_total{app=\"order-service\"}[5m]))\n > 0.05\n for: 3m\n labels:\n severity: critical\n annotations:\n summary: \"High 5xx error rate on order-service\"\n description: \"Error rate is {{ $value | humanizePercentage }} over the last 5 minutes\"\n - alert: HighLatency\n expr: |\n histogram_quantile(0.99,\n sum(rate(http_request_duration_seconds_bucket{app=\"order-service\"}[5m])) by (le)\n ) > 2\n for: 5m\n labels:\n severity: warning\n annotations:\n summary: \"P99 latency exceeding 2s on order-service\"\n - alert: PodRestartsFrequent\n expr: |\n increase(kube_pod_container_status_restarts_total{namespace=\"production\", container=\"order\"}[1h]) > 3\n for: 0m\n labels:\n severity: critical\n annotations:\n summary: \"Pod {{ $labels.pod }} restarting frequently\"\n\"",
|
||
"language": "yaml",
|
||
"explanation": "本题考察 Prometheus 监控体系的配置,包括 ServiceMonitor 的自动发现机制和 PromQL 告警规则。需要理解 PromQL 的聚合函数、直方图百分位计算以及告警条件的含义。",
|
||
"source": null,
|
||
"related": [],
|
||
"sub_questions": [
|
||
{
|
||
"index": 1,
|
||
"type": "short_answer",
|
||
"question": "HighErrorRate 告警规则中的 PromQL 表达式计算的是什么指标?当该值大于 0.05 且持续 3 分钟时,说明什么问题?",
|
||
"answer": "该表达式计算的是 HTTP 5xx 错误率(5xx 请求数 / 总请求数)。大于 0.05 意味着超过 5% 的请求返回了服务器错误,持续 3 分钟说明这不是瞬时毛刺而是持续性问题,触发 critical 级别告警。",
|
||
"keywords": [
|
||
"5xx 错误率",
|
||
"5%",
|
||
"HTTP 5xx",
|
||
"持续性",
|
||
"critical"
|
||
],
|
||
"scoring_rubric": "答出错误率含义(1分)、5% 阈值(1分)、持续时间的过滤意义(1分),共3分。",
|
||
"explanation": "sum(rate(http_requests_total{status=~\"5..\"}[5m])) 计算 5 分钟内 5xx 状态码的请求速率,除以总请求速率得到错误率。rate() 使用 5 分钟窗口平滑瞬时波动。for: 3m 表示条件必须连续满足 3 分钟才触发告警,过滤掉短暂的网络抖动或重启导致的瞬时错误。如果错误率持续 >5%,通常意味着代码 bug、依赖服务故障或资源不足。"
|
||
},
|
||
{
|
||
"index": 2,
|
||
"type": "single_choice",
|
||
"question": "HighLatency 告警使用了 histogram_quantile(0.99, ...) 表达式,这表示什么含义?如果 P99 延迟超过 2 秒,说明什么?",
|
||
"options": {
|
||
"A": "所有请求中有 1% 的请求延迟超过 2 秒",
|
||
"B": "所有请求中有 99% 的请求延迟超过 2 秒",
|
||
"C": "平均延迟超过 2 秒",
|
||
"D": "最大延迟超过 2 秒"
|
||
},
|
||
"answer": "A",
|
||
"explanation": "histogram_quantile(0.99, ...) 计算的是第 99 百分位延迟值,即 99% 的请求延迟低于该值,只有 1% 的请求延迟超过该值。当 P99 > 2 秒时,意味着长尾延迟较高——虽然大多数请求响应正常,但有 1% 的请求耗时超过 2 秒。这在用户体验上可能表现为部分用户遇到明显卡顿。P99 是比平均值更能反映尾部延迟的关键指标。"
|
||
},
|
||
{
|
||
"index": 3,
|
||
"type": "single_choice",
|
||
"question": "PodRestartsFrequent 告警使用 kube_pod_container_status_restarts_total 指标配合 increase 函数。以下哪个说法正确?",
|
||
"options": {
|
||
"A": "restarts_total 是 Counter 类型,increase 统计过去 1 小时内重启次数增量",
|
||
"B": "restarts_total 是 Gauge 类型,increase 统计过去 1 小时的最大值",
|
||
"C": "restarts_total 是 Histogram 类型,increase 统计分布变化",
|
||
"D": "increase 函数只能用于计算 CPU 使用量"
|
||
},
|
||
"answer": "A",
|
||
"explanation": "kube_pod_container_status_restarts_total 是 Kubernetes 暴露的 Counter 类型指标,值只会递增不会递减。increase(kube_pod_container_status_restarts_total[1h]) > 3 表示过去 1 小时内该容器重启次数超过 3 次。for: 0m 表示一旦满足条件立即触发告警(不需要等待持续评估)。频繁重启通常是 livenessProbe 失败、OOMKilled 或应用崩溃的信号,需要立即关注。"
|
||
}
|
||
]
|
||
},
|
||
{
|
||
"id": "cr-005",
|
||
"type": "code_reading",
|
||
"difficulty": 4,
|
||
"tags": [
|
||
"opentelemetry",
|
||
"kubernetes",
|
||
"prometheus",
|
||
"loki"
|
||
],
|
||
"question": "以下是一个 OpenTelemetry Collector 的部署配置,阅读后回答问题。",
|
||
"code": "apiVersion: v1\nkind: ConfigMap\nmetadata:\n name: otel-collector-config\n namespace: observability\ndata:\n config.yaml: |\n receivers:\n otlp:\n protocols:\n grpc:\n endpoint: 0.0.0.0:4317\n http:\n endpoint: 0.0.0.0:4318\n prometheus:\n config:\n scrape_configs:\n - job_name: 'kubernetes-pods'\n kubernetes_sd_configs:\n - role: pod\n relabel_configs:\n - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]\n action: keep\n regex: true\n processors:\n batch:\n timeout: 5s\n send_batch_size: 1000\n memory_limiter:\n check_interval: 1s\n limit_mib: 512\n spike_limit_mib: 128\n attributes:\n actions:\n - key: environment\n action: upsert\n value: production\n exporters:\n otlp/traces:\n endpoint: jaeger-collector.observability:4317\n tls:\n insecure: false\n prometheus:\n endpoint: 0.0.0.0:8889\n namespace: otel\n loki:\n endpoint: http://loki-gateway.observability:3100/loki/api/v1/push\n service:\n pipelines:\n traces:\n receivers: [otlp]\n processors: [memory_limiter, batch]\n exporters: [otlp/traces]\n metrics:\n receivers: [otlp, prometheus]\n processors: [attributes, batch]\n exporters: [prometheus]\n logs:\n receivers: [otlp]\n processors: [memory_limiter, batch]\n exporters: [loki]",
|
||
"language": "yaml",
|
||
"explanation": "本题考察 OpenTelemetry Collector 的完整配置,包括接收器、处理器、导出器和管道的组织方式。需要理解 OTel Collector 的数据流架构和各组件的作用。",
|
||
"source": null,
|
||
"related": [],
|
||
"sub_questions": [
|
||
{
|
||
"index": 1,
|
||
"type": "single_choice",
|
||
"question": "在 traces pipeline 中,processors 的配置顺序是 [memory_limiter, batch]。memory_limiter 放在 batch 之前的原因是什么?",
|
||
"options": {
|
||
"A": "为了先对数据进行压缩再批处理",
|
||
"B": "为了在批处理之前检查内存使用量,在内存超限时拒绝接收新数据,防止 OOM 崩溃",
|
||
"C": "memory_limiter 必须始终在 batch 之前,这是 OTel Collector 的硬性要求",
|
||
"D": "为了先过滤不需要的 trace 再发送"
|
||
},
|
||
"answer": "B",
|
||
"explanation": "memory_limiter 是 OTel Collector 的保护性处理器,它在每 1 秒(check_interval)检查进程内存使用量。当内存超过 limit_mib(512MB)时,它会触发垃圾回收;当内存超过 limit_mib + spike_limit_mib(640MB)时,会拒绝接收新数据并返回 429 错误。将其放在 batch 之前,是因为 batch 处理会批量积累数据,如果不先进行内存检查,可能导致内存持续增长直至 OOM。这是 OTel Collector 的最佳实践配置顺序:memory_limiter → batch。"
|
||
},
|
||
{
|
||
"index": 2,
|
||
"type": "short_answer",
|
||
"question": "该 OTel Collector 配置了三个 pipeline(traces、metrics、logs),请简述每条 pipeline 的数据流向,包括接收器、处理器和导出器分别是什么?",
|
||
"answer": "traces pipeline: OTLP → memory_limiter → batch → Jaeger (OTLP/gRPC)。metrics pipeline: OTLP + Prometheus (auto-discovery) → attributes (注入 environment 标签) → batch → Prometheus exporter (暴露 :8889 供 Prometheus 抓取)。logs pipeline: OTLP → memory_limiter → batch → Loki (HTTP push)。",
|
||
"keywords": [
|
||
"traces",
|
||
"metrics",
|
||
"logs",
|
||
"OTLP",
|
||
"Prometheus",
|
||
"Loki",
|
||
"Jaeger"
|
||
],
|
||
"scoring_rubric": "traces 流向正确(1分)、metrics 流向正确含 attributes 处理(1分)、logs 流向正确(1分),共3分。",
|
||
"explanation": "OTel Collector 的核心架构是 Receiver → Processor → Exporter 的管道模式。traces pipeline 通过 OTLP 接收链路追踪数据,经过内存保护和批处理后导出到 Jaeger。metrics pipeline 同时从 OTLP 和 Prometheus 自动发现接收指标数据,通过 attributes 处理器为所有指标注入 environment=production 标签,再通过 Prometheus exporter 暴露在 :8889 端口供 Prometheus 抓取。logs pipeline 通过 OTLP 接收日志数据,经保护和批处理后推送到 Loki。三条 pipeline 共享 receivers(如 OTLP),但各自的 processor 和 exporter 配置独立。"
|
||
},
|
||
{
|
||
"index": 3,
|
||
"type": "single_choice",
|
||
"question": "prometheus receiver 配置中的 kubernetes_sd_configs role=pod 和 relabel_configs 用于什么目的?",
|
||
"options": {
|
||
"A": "从 Kubernetes API 发现 Pod 并根据 prometheus.io/scrape 注解决定是否抓取该 Pod 的 metrics",
|
||
"B": "将 Pod 标签同步到 Prometheus 的服务发现配置",
|
||
"C": "自动为所有 Pod 安装 Prometheus exporter",
|
||
"D": "从 Kubernetes API 发现 Service 并抓取 Service 的 metrics"
|
||
},
|
||
"answer": "A",
|
||
"explanation": "kubernetes_sd_configs 配置 role=pod 表示 OTel Collector 会通过 Kubernetes API 自动发现集群中的所有 Pod。relabel_configs 中的 action=keep 配合 regex: true 表示只有当 Pod 带有 prometheus.io/scrape=true 注解时,才会被纳入抓取目标。这是云原生环境下 Prometheus 动态服务发现的标准模式——应用只需添加注解即可自动被监控,无需修改 Prometheus 配置。这种零侵入式的监控集成方式在微服务架构中尤为重要。"
|
||
}
|
||
]
|
||
}
|
||
]
|
||
} |