This repository has been archived on 2026-05-24. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
all-in-kingsoft/hhs/gRPC/1. Protobuf 基础篇/03-字段编号与前向兼容.md
T
2026-05-11 19:02:38 +08:00

307 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
tags: [gRPC, Protobuf, field number, reserved, backward compatibility, forward compatibility, versioning]
create time: 2026-05-11 16:40
---
# 字段编号与前向兼容
## 概述
每一个字段都有一个 tag number,这行简单的数字背后藏着 Protobuf 最核心的设计原则:**向后兼容**。理解这套机制,你就能放心地修改 protobuf schema 而不用担心打坏线上服务。
> [!note] 核心概念速记
> - **向后兼容**(Old → New):旧版客户端跑新版服务端返回的数据 — Protobuf 保证**未知字段被安全忽略**
> - **向前兼容**(New → Old):新版客户端跑旧版服务端返回的数据 — 缺失字段取**类型默认值**
> [!warning] 注意
> 一旦字段编号被分配并部署,就**永远不能再复用**它。这是 Protobuf 的硬伤——编号就像 UUID,一旦发出去就是它的了。
## Tag Number 分配规则
每个字段的 tag number(通常称为 field number)取值范围为 **1 ~ 536,870,911**(即 `2^29 - 1`)。这个范围不是随机的:
```protobuf
message User {
string id = 1; // 核心标识符,高频使用 → 留给 1~15
string name = 2;
string email = 3;
int32 age = 4;
bool active = 5;
string phone = 6;
string address = 7;
Role role = 8;
// ... 中间跳过一些编号供未来添加 ...
// (reserved 10 to 20)
google.protobuf.Timestamp created_at = 100;
google.protobuf.Timestamp updated_at = 101;
repeated string tags = 102;
}
```
### Wire Encoding 原理解析
Protobuf 使用 **varint encoding**(变长整数编码),tag number 越小占用的字节越少:
| 编号范围 | Wire Encoding 大小 | 说明 |
|---------|-------------------|------|
| 1 ~ 15 | 1 byte | 黄金区间,预留给你的核心高频字段 |
| 16 ~ 2047 | 2 bytes | 次优先区间 |
| 2048+ | 3~5 bytes | 低频字段可放这里 |
```mermaid
flowchart LR
subgraph F1["Field Number = 3"]
A1["field_number = 3"] --> B1["<< 3"]
B1 --> C1["24"]
C1 --> D1["| wire_type 0"]
D1 --> E1["tag = 24 = 0x18"]
E1 --> F1["1 byte ✅"]
end
subgraph F2["Field Number = 100"]
A2["field_number = 100"] --> B2["<< 3"]
B2 --> C2["800"]
C2 --> D2["| wire_type 0"]
D2 --> E2["tag = 800 = 0x320"]
E2 --> F2["2 bytes ⚠️"]
end
style F1 fill:#d4edda
style F2 fill:#fff3cd
```
> [!example] 公式
> `tag = (field_number << 3) | wire_type`
> - `<< 3` 等价于 `field_number × 8`,把高 5 位留给 field number
> - 低 3 位存放 wire type(0=varint, 1=64-bit, 2=length-delimited, ...)
对于高频通信的消息体,节省 1 byte *per message* × 每秒百万调用 = 可观的带宽节省。这就是为什么建议 core fields 用 1~15。
### 合理的 Field Numbering 策略
```protobuf
// user/v1/user.proto
message User {
// === Core fields (1-9): 核心字段,几乎每次都会序列化 ===
string id = 1;
string name = 2;
string email = 3;
// === Secondary fields (10-19): 常用但非必需 ===
string phone = 10;
string avatar_url = 11;
Role role = 12;
bool active = 13;
// === Tertiary fields (20-99): 偶尔使用 ===
string bio = 20;
string website = 21;
Location location = 22;
// === Audit & metadata (100-199): 系统字段,低频 ===
google.protobuf.Timestamp created_at = 100;
google.protobuf.Timestamp updated_at = 101;
string created_by = 102;
string updated_by = 103;
// === Feature flags / experimental (900-999): 灰度测试用 ===
bool new_ui_enabled = 900;
}
```
> [!tip] 预留块的好处
> 如果你的 User 消息已经有 13 个字段,未来需要新增 5 个字段,你只需要在 10~19 之间找空位。如果所有字段从 1 开始连续排列,每加一个都需要改后面所有的编号——而且已经部署的旧客户端会认为新编号的字段属于不同的语义。
## Reserved 保留字段
当你删除或重命名字段时,必须用 `reserved` 声明来防止后人误用相同的编号:
```protobuf
message User {
// ---- 当前活跃字段 ----
string id = 1;
string name = 2;
string email = 3;
int32 age = 4;
// ---- 已废弃字段的编号保留 ----
reserved "mobile"; // 之前叫 mobile 的字段已删除
reserved 5, 6; // 编号 5 和 6 已释放,禁止复用
reserved 7 to 10; // 编号 7~10 连续保留
}
```
### 删除字段的正确姿势
三步走,确保平滑过渡:
```mermaid
flowchart TD
S1["📝 步骤 1: 用 reserved 占位"] --> S2["string old_field = 5;\n→\nreserved 5;"]
S2 --> S3["📢 步骤 2: 协调消费方迁移"]
S3 --> S4["(版本发布窗口内切换)"]
S4 --> S5["✅ 步骤 3: 下次编译锁定\n有人复用 → compile error"]
S5 --> Safe["后续正式移除"]
style S2 fill:#fff3cd
style S4 fill:#fff3cd
style Safe fill:#d4edda
```
> [!question] 为什么不能只删字段不 reserved?
> 如果没有 reserved,同事新建字段时使用相同编号:`string feature_flag = 5;`。旧版本客户端读到这个字节流时,会把 feature_flag 的值当成旧版 mobile 字段的值——数据语义完全错位,bug 极难排查。
## 兼容性矩阵(重点章节)
Protobuf 的设计确保了大部分 schema 变更不会破坏现有二进制协议:
| 操作 | 向后兼容? | 向前兼容? | 说明 |
|------|-----------|-----------|------|
| 新增字段 | ✅ | ✅ | 老客户端忽略未知编号;新客户端用默认值 |
| 删除字段 | ✅ | ✅ | 老客户端读取已有数据;新客户端用默认值 |
| 修改字段类型 | ❌ | ❌ | 新旧对同一编号解读不同 |
| 修改字段编号 | ❌ | ❌ | 同编号对应不同语义 |
| 修改枚举值名称 | ✅ | ✅ | 枚举值名不影响 wire format(传输的是数值) |
| 新增枚举值 | ✅ | ⚠️ | 旧客户端收到未知枚举值回退为 0(首个值) |
| 删除枚举值 | ❌ | ⚠️ | 旧客户端收到未知枚举值回退为 0 |
| 单个 repeated 改为 non-repeated | ⚠️ | ❌ | 有数据的单元素列表可互转,多元素场景不兼容 |
### 实战:安全地扩展消息
假设你有一个在线上运行的 v1 proto:
```protobuf
// v1 - 当前生产版本
message GetUserResponse {
string id = 1;
string name = 2;
string email = 3;
}
```
**需求:添加 `avatar_url` 和 `role` 两个字段,同时删除 `email`。**
❌ **错误示范 — 直接复用已被 reserved 的编号:**
```protobuf
// v2 - ❌ 编译失败
message GetUserResponse {
string id = 1;
string name = 2;
reserved 3; // email 的编号已保留
string avatar_url = 3; // ← 编译报错:field number 3 is reserved
}
```
✅ **正确做法 — 使用新编号 + reserved 占位:**
```protobuf
// v2 - ✅ 安全演进
message GetUserResponse {
string id = 1;
string name = 2;
reserved 3; // 原 email 编号锁定,防止后人误用
string avatar_url = 4; // 新字段分配新编号
User_Role role = 5;
}
```
> [!note] 正确的分步迁移方案
> 1. 先加字段 `avatar_url = 4`、`role = 5`,email 继续保留。
> 2. 服务端双写:同时返回 email 和 avatar_url。
> 3. 客户端升级,切换到使用 avatar_url。
> 4. 确认旧客户端已淘汰后,标记 email 为 reserved,下次发版正式移除。
## 版本演进策略
对于大型项目,建议使用 package-level versioning 来管理 schema 演进:
### 目录结构与命名约定
```
protos/
├── user/
│ └── v1/
│ ├── user.proto
│ ├── auth.proto
│ └── error.proto
└── order/
└── v1/
├── order.proto
└── payment.proto
```
```protobuf
// option go_package 包含版本路径
option go_package = "github.com/example/service/user/v1;userpb";
// import 路径与目录结构一致
import "user/v1/user.proto";
```
### Major Version 迁移方案
当需要做不兼容变更时(如改字段类型、重构消息结构):
```mermaid
flowchart LR
subgraph A["方案 A: v2 独立演进 🌟 推荐"]
direction TB
A1["user/v1/user.proto"] --> A2["新旧并存\nGateway 层做转换"]
A3["user/v2/user.proto"] --> A2
A2 --> A4["迁移完成\n停用 v1"]
end
subgraph B["方案 B: 原地破坏 ❌ 高风险"]
B1["user/v1/user.proto\n直接改"] --> B2["已部署端受影响"]
end
style A fill:#d4edda
style B fill:#f8d7da
style A4 fill:#28a745,color:#fff
style B2 fill:#dc3545,color:#fff
```
| 维度 | 方案 A(v2 独立文件) | 方案 B(原地修改) |
|------|---------------------|-------------------|
| 风险等级 | 低 | **极高** |
| 线上影响 | Gateway 透明转换 | 所有端同时断裂 |
| 回滚成本 | 切回 v1 即可 | 几乎无法回滚 |
| 适用场景 | 所有已发布服务 | 仅限内部未发布 proto |
> [!tip] 灰度策略:双字段过渡法
> 如果需要在同一消息中过渡一个新字段到旧字段,分三阶段进行:
> ```protobuf
> // 阶段一: 服务端双写两个字段
> message User {
> string legacy_name = 1; // 旧字段,逐步弃用
> string display_name = 2; // 新字段,逐步启用
> }
>
> // 阶段二: 客户端优先读 display_name, 回退到 legacy_name
> //
> // 阶段三: 确认全部升级后移除 legacy_name (步骤见上方「删除字段的正确姿势」)
> ```
## 最佳实践
- **为每个 microservice 预留独立 namespace**:`package service_name.version`。
- **不要重复使用 field numbers**:即使在同一个文件中删除了字段也要 reserved。
- **核心高频字段编号保持在 1~15**:节省 wire 编码开销。
- **重大变更走 v2 而不是改现有文件**:降低线上风险。
- **在 CI 中加入 proto linter**(如 buf lint):自动化检查编号冲突和命名规范。
## 关联笔记
- [[hhs/gRPC/1. Protobuf 基础篇/01-Protobuf 语法与消息定义]] — Protobuf 消息定义基础语法
- [[hhs/gRPC/1. Protobuf 基础篇/02-数据类型详解]] — 字段可用的所有数据类型及默认值规则
- [[hhs/gRPC/1. Protobuf 基础篇/04-Oneof 与包装类型]] — Oneof 的 field numbering 有特殊规则
- [[hhs/gRPC/6. 工程实践篇/18-模块拆分与 proto 规范]] — Proto 文件的工程组织与命名规范