Files
cs-note/hhs/MS/04-可观测性/01-Metrics监控.md
T
2026-05-24 11:42:38 +08:00

5.8 KiB

tags, create time
tags create time
microservice
metrics
prometheus
grafana
sre
2026-05-05

Metrics 监控

概述

Metrics 回答的问题是:系统现在健康吗?——通过聚合后的数值揭示趋势。

graph TB
    App["应用服务"] -->|Push / Scrape| PROM[(Prometheus)]
    PROM --> GRAF["Grafana Dashboard"]
    PROM --> ALERT["Alertmanager"]
    
    style PROM fill:#e3f2fd
    style GRAF fill:#fff3e0
    style ALERT fill:#fce4ec

指标类型

类型 含义 特点 示例
Counter 只增不减的计数器 可 Reset(重启) http_requests_total
Gauge 可升可降的仪表盘 反映当前状态 queue_depth, cpu_temp
Histogram 样本分布,自动分桶 计算 P50/P90/P99 api_latency_seconds
Summary 类似 Histogram,客户端算分位 Go SDK 默认类型 grpc_duration_seconds

Counter vs Gauge 场景选择

flowchart LR
    Q{"这个值是<br/>只增不减的吗?"}
    
    Q -- "是" --> C["Counter<br/>请求数、错误数、订单量"]
    Q -- "否" --> G["Gauge<br/>在线用户数、队列长度、内存使用量"]
    
    style C fill:#c8e6c9
    style G fill:#bbdefb

RED 方法 (针对有状态服务)

适用于 API、微服务等有明确请求/响应的服务。

指标 公式 说明
Rate rate(http_requests_total[5m]) 每秒请求量 (QPS)
Errors rate(http_requests_total{status="5xx"}[5m]) 每秒错误数
Duration histogram_quantile(0.99, rate(api_latency_bucket[5m])) P99 响应时间

[!tip] PromQL 关键函数

  • rate() — 计算 Counter 每秒增长率(必须用于 Counter)
  • irate() — 即时速率,对突发更敏感
  • histogram_quantile(0.99, ...) — 从直方图计算分位数
  • increase() — 时间段内的增长总量
  • avg() / max() / min() — 基础聚合函数

USE 方法 (针对基础设施)

适用于 CPU、内存、网络、磁盘等底层资源监控。

指标 说明 Grafana PromQL 示例
Utilization 使用率 1 - avg(rate(node_cpu_seconds_total{mode="idle"}[5m]))
Saturation 饱和度 node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes
Errors 错误数 rate(node_network_receive_errs_total[5m])

Go Prometheus 集成

var (
    httpRequestsTotal = prometheus.NewCounterVec(
        prometheus.CounterOpts{
            Name: "http_requests_total",
            Help: "Total HTTP requests by method and status",
        },
        []string{"method", "status"},
    )
    
    apiLatency = prometheus.NewHistogramVec(
        prometheus.HistogramOpts{
            Name:    "api_latency_seconds",
            Help:    "API latency distribution",
            Buckets: prometheus.DefBuckets, // [0.005, 0.01, ..., 10]
        },
        []string{"endpoint"},
    )
)

func init() {
    prometheus.MustRegister(httpRequestsTotal, apiLatency)
}

// Middleware 中使用
func MetricsMiddleware(next http.Handler) http.Handler {
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        start := time.Now()
        next.ServeHTTP(w, r)
        
        duration := time.Since(start).Seconds()
        httpRequestsTotal.WithLabelValues(r.Method, fmt.Sprintf("%d", w.Status())).Inc()
        apiLatency.WithLabelValues(r.URL.Path).Observe(duration)
    })
}

Dashboard 设计原则

一个优秀的 Dashboard 应该让任何团队成员在 30 秒内了解服务的整体状态:

flowchart TB
    subgraph DASH["Service Health Dashboard"]
        ROW1["🔴 可用性 & 错误率 — 第一眼判断"]
        ROW2["🟡 性能指标 — P50/P90/P99 趋势"]
        ROW3["🔵 基础设施 — CPU/内存/连接数"]
        ROW4["⚫ 业务指标 — 订单量/支付成功率"]
    end
    
    ROW1 --> JUDGE{是否异常?}
    ROW2 --> JUDGE
    ROW3 --> JUDGE
    ROW4 --> JUDGE
    
    JUDGE --"否" --> NORMAL["一切正常 ✓"]
    JUDGE --"是" --> ALERT["触发告警 → On-Call"]

Dashboard 布局模板

┌─────────────────────────────────────────────────┐
│  Row 1: 🔴 Service Availability                  │
│  ├─ QPS (Rate)          ┌─ Error Rate (%)       │
│  ├─ Active Connections   └─ 5xx Count           │
├─────────────────────────────────────────────────┤
│  Row 2: 🟡 Performance                          │
│  ├─ P50 Latency         ┌─ P90 Latency          │
│  ├─ P99 Latency         └─ Slow Requests (>1s)  │
├─────────────────────────────────────────────────┤
│  Row 3: 🔵 Infrastructure                      │
│  ├─ CPU %               ┌─ Memory Usage         │
│  ├─ GC Pause Time       └─ Goroutine Count     │
├─────────────────────────────────────────────────┤
│  Row 4: ⚫ Business Metrics                     │
│  ├─ Orders/Minute       ┌─ Payment Success Rate │
│  └─ New Users/Day       └─ Failed Transactions  │
└─────────────────────────────────────────────────┘

关联笔记