Files
cs-note/hzh/Gen2D/07-可观测性.md
T

196 lines
7.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
tags: [observability, prometheus, grafana, metrics, alerting, monitoring]
create time: 2026-06-03 10:30
---
# 07. 可观测性
## 概述
35 个 Prometheus 指标 + 3 个 Grafana 仪表盘 + 10 条告警规则,覆盖全栈,让系统运行状态一目了然。
---
## 正文
## 指标体系总览
Gen2D 遵循 Prometheus 命名最佳实践,所有指标使用 `gen2d_` 前缀,共 **35 个指标**,分为 **6 大组**:
| 组 | 指标数 | 采集方式 | 侵入性 |
|:--:|:------:|---------|:------:|
| HTTP 层 | 5 | Gin 中间件自动采集 | 零侵入 |
| 限流层 | 2 | 限流中间件自动采集 | 零侵入 |
| 队列层 | 5 | 队列实现内部埋点 | 低 |
| 协程池层 | 6 | 池内部埋点 | 低 |
| Pipeline 层 | 5 | 管线节点回调 | 低 |
| 基础设施层 | 5 | 连接状态监控 | 低 |
---
## 五大指标组详解
### HTTP 层指标
Gin 中间件自动采集,**零业务代码侵入**。
| 指标名 | 类型 | 标签 | 说明 |
|--------|------|------|------|
| `gen2d_http_requests_total` | Counter | method, path, status | 请求总量 |
| `gen2d_http_request_duration_seconds` | Histogram | method, path | 请求延迟分布 |
| `gen2d_http_request_size_bytes` | Histogram | method, path | 请求体大小 |
| `gen2d_http_response_size_bytes` | Histogram | method, path | 响应体大小 |
| `gen2d_http_requests_in_flight` | Gauge | — | 当前并发请求数 |
> [!tip] FullPath() 的关键作用
> 使用路由模板 `/api/v1/tasks/:taskId` 而非实际路径 `/api/v1/tasks/abc123`,避免高基数标签导致 Prometheus 内存爆炸。
### 限流层指标
| 指标名 | 类型 | 标签 | 说明 |
|--------|------|------|------|
| `gen2d_ratelimit_requests_total` | Counter | scope, endpoint, result | 限流决策总量 |
| `gen2d_ratelimit_remaining_tokens` | Gauge | scope, endpoint | 剩余令牌数 |
- `scope`:`user` / `global`
- `result`:`allowed` / `denied`
### 任务队列层指标
| 指标名 | 类型 | 标签 | 说明 |
|--------|------|------|------|
| `gen2d_queue_depth` | Gauge | driver | 当前队列积压深度 |
| `gen2d_queue_submitted_total` | Counter | driver | 入队总量 |
| `gen2d_queue_consumed_total` | Counter | driver | 出队总量 |
| `gen2d_queue_submit_duration_seconds` | Histogram | driver | 入队耗时 |
| `gen2d_queue_errors_total` | Counter | driver, error_type | 队列错误总量 |
- `driver`:`memory` / `rabbitmq`
### 协程池层指标
| 指标名 | 类型 | 标签 | 说明 |
|--------|------|------|------|
| `gen2d_pool_active_workers` | Gauge | — | 活跃 worker 数 |
| `gen2d_pool_queued_tasks` | Gauge | — | 池内排队任务数 |
| `gen2d_pool_submitted_total` | Counter | — | 提交到池的任务总量 |
| `gen2d_pool_completed_total` | CounterVec | result | 完成的任务总量 |
| `gen2d_pool_rejected_total` | Counter | — | 被拒绝的任务 |
| `gen2d_pool_task_duration_seconds` | Histogram | — | 任务执行耗时 |
### Pipeline 业务层指标
| 指标名 | 类型 | 标签 | 说明 |
|--------|------|------|------|
| `gen2d_pipeline_total` | CounterVec | status | Pipeline 执行总量 |
| `gen2d_pipeline_duration_seconds` | HistogramVec | status | 端到端耗时 |
| `gen2d_pipeline_stage_duration_seconds` | HistogramVec | stage | 各阶段耗时 |
| `gen2d_pipeline_retries_total` | CounterVec | stage | 各阶段重试次数 |
| `gen2d_pipeline_tasks_active` | Gauge | — | 当前执行中的 Pipeline 数 |
> [!tip] stage_duration 定位瓶颈
> 通过 `stage` 标签(如 `asset_generator`、`quality_check`)可以精确定位哪个阶段是性能瓶颈。
### 基础设施层指标
| 指标名 | 类型 | 标签 | 说明 |
|--------|------|------|------|
| `gen2d_redis_operations_total` | CounterVec | op, result | Redis 操作总量 |
| `gen2d_redis_operation_duration_seconds` | HistogramVec | op | Redis 操作延迟 |
| `gen2d_redis_connection_pool_size` | Gauge | — | Redis 连接池大小 |
| `gen2d_rabbitmq_connection_status` | Gauge | — | RabbitMQ 连接状态 |
| `gen2d_rabbitmq_reconnect_total` | Counter | — | RabbitMQ 重连次数 |
---
## 中间件集成
### Metrics 中间件工作流程
```go
func Metrics() gin.HandlerFunc {
return func(c *gin.Context) {
path := c.FullPath() // 路由模板,避免高基数
HTTPRequestsInFlight.Inc() // 进入时 +1
defer HTTPRequestsInFlight.Dec() // 离开时 -1
reqSize := c.Request.ContentLength // 请求体大小 before
c.Next() // 执行后续链
elapsed := time.Since(start) // 耗时 after
HTTPRequestsTotal.WithLabelValues(...).Inc()
}
}
```
**关键设计**:
- `in_flight` 使用 `defer` 保证异常退出也能正确递减
- 请求体大小在 `c.Next()` 之前采集(此时 Content-Length 已知)
- 响应体大小在 `c.Next()` 之后采集(此时 Writer 已写入)
---
## Grafana 仪表盘
| 仪表盘 | 用途 | 关键面板 |
|--------|------|---------|
| **Overview** | 全局概览 | 请求量、错误率、延迟 P50/P95/P99、活跃连接 |
| **Pipeline** | 管线监控 | 成功率、各阶段耗时、重试率、活跃任务数 |
| **Infrastructure** | 基础设施 | Redis/RabbitMQ 状态、队列深度、池饱和度 |
---
## 告警规则
共 **10 条告警规则**,覆盖限流、队列、协程池、管线和基础设施:
| 告警名 | 级别 | 条件 | 说明 |
|--------|:----:|------|------|
| `RateLimitHighDenialRate` | ⚠️ | 限流拒绝率 > 5%(持续 5m) | 可能遭受攻击或配置过严 |
| `QueueBacklog` | ⚠️ | 队列深度 > 50(持续 2m) | 消费能力不足 |
| `QueueErrors` | 🔴 | 5 分钟内错误 > 5 次 | 队列服务异常 |
| `PoolSaturation` | ⚠️ | 活跃 worker 占比 > 90%(持续 5m) | 考虑扩容 |
| `PoolTaskRejected` | ⚠️ | 5 分钟内拒绝 > 5 个 | 池容量不足 |
| `PipelineSuccessRateLow` | 🔴 | 成功率 < 90%(持续 10m) | 生成服务异常 |
| `RedisDown` | 🔴 | 连接池大小 = 0(持续 1m) | Redis 不可用 |
| `RabbitMQDisconnected` | 🔴 | 连接状态 = 0(持续 1m) | RabbitMQ 断连 |
| `HighErrorRate` | 🔴 | 5xx 错误率 > 5%(持续 5m) | 服务异常 |
| `HighLatency` | ⚠️ | P95 延迟 > 5s(持续 5m) | 影响用户体验 |
> [!note] 告警级别说明
> Critical 表示需要立即处理,Warning 表示需要关注但不紧急。
---
## 基础设施指标
除业务指标外,Gen2D 还监控外部依赖的健康状态:
```mermaid
flowchart LR
subgraph "基础设施监控"
R["Redis"] -->|"operations_total"| M["Prometheus"]
R -->|"connection_pool_size"| M
Q["RabbitMQ"] -->|"connection_status"| M
Q -->|"reconnect_total"| M
end
M --> G["Grafana"]
M --> A["AlertManager"]
style R fill:#fce4ec,stroke:#c62828
style Q fill:#fff3e0,stroke:#f57c00
style M fill:#e8f5e9,stroke:#388e3c
style G fill:#e8eaf6,stroke:#303f9f
style A fill:#ffebee,stroke:#b71c1c
```
- **Redis**:操作延迟、成功率、连接池大小
- **RabbitMQ**:连接状态(1=connected, 0=disconnected)、重连次数
---
## 关联文档
- [[10-中间件链]] — Metrics 中间件的挂载位置
- [[09-限流]] — 限流指标的采集方式
- [[14-部署架构]] — Prometheus + Grafana 的部署配置