This repository has been archived on 2026-05-24. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
all-in-kingsoft/hhs/gRPC/1. Protobuf 基础篇/02-数据类型详解.md
T
2026-05-11 19:02:38 +08:00

419 lines
16 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
tags: [gRPC, Protobuf, scalar types, wrapper types, WKT, repeated, packed, oneof, map, wire encoding]
create time: 2026-05-11 16:40
---
# 数据类型详解
## 概述
Protobuf 的类型系统看起来简单,但有很多容易被忽略的细节:`optional`/`required` 的区别、packed vs unpacked repeated 编码差异、以及 Well-Known Types 的威力。这篇帮你把常见坑一次性踩完。
> [!question] 为什么 Protobuf 没有 required?
> 早期的 proto3 移除了 `required`/`optional` 关键字,因为工程实践中很难真正验证——服务端删除了字段后,客户端无法区分"字段没传"和"服务端没设值"。**如果需要保证某个字段一定存在,该用什么方式替代?** 提示:见文末 `oneof` 用法。
## Protobuf 类型体系一览
在深入每个类型之前,先看全貌:
```mermaid
graph TD
A["Protobuf 类型系统"] --> B["Scalar Types\n标量类型"]
A --> C["Composite Types\n复合类型"]
A --> D["Well-Known Types\n内置类型"]
B --> B1["整数系: int32 / int64 / uint32 / uint64 / sint32 / sint64"]
B --> B2["浮点系: float / double"]
B --> B3["其他: bool / string / bytes"]
C --> C1["repeated\n(动态列表)"]
C --> C2["map<string, T>\n(键值对)"]
C --> C3["message\n(自定义结构)"]
C --> C4["oneof\n(互斥字段)"]
D --> D1["Timestamp\n(time.Time)"]
D --> D2["Duration\n(time.Duration)"]
D --> D3["StringValue\n(*string 指针)"]
D --> D4["Any / Value / Struct\n(通用 JSON)"]
D --> D5["FieldMask\n(partial update)"]
```
### proto2 vs proto3 关键差异
| 特性 | proto2 | proto3 |
|------|--------|--------|
| `required` / `optional` | 支持 | ❌ 移除(proto3 用默认零值语义) |
| `enum default value` | 不允许 0 以外的默认值 | ✅ 允许任意枚举值作为默认 |
| map | ❌ 不支持 | ✅ 原生支持 |
| repeated packed | 需显式声明 `[packed = true]` | ✅ 数字类型默认 packed |
| `Has()` 判断 | 自动生成 | ❌ 不再为 scalar 生成(wrapper type 替代) |
> [!tip] proto3 的 optional 回来了!
> 虽然 proto3 最初去掉了 optional,但从 **protobuf 3.12+** 开始重新引入了 `optional` 关键字,不过它仍然受 wire compatibility 限制——加上 optional 后会改变 field number 的行为,所以生产环境中更推荐用 **wrapper types**。
## Scalar Types 标量类型
Protobuf 提供了一套语言无关的标量类型,每种都有确定的 wire encoding。选型的核心原则是:**在保证正确性的前提下,选最小的类型**。
| Protobuf 类型 | Go 生成类型 | Wire Encoding | 说明 |
|---------------|------------|---------------|------|
| `double` | `float64` | 8 bytes | 双精度浮点 |
| `float` | `float32` | 4 bytes | 单精度浮点 |
| `int32` | `int32` | varint | **最常用**,小整数高效编码 |
| `int64` | `int64` | zigzag varint | 大整数或时间戳 |
| `uint32` | `uint32` | varint | 无符号 32 位 |
| `uint64` | `uint64` | varint | 无符号 64 位 |
| `sint32` | `int32` | zigzag varint | 有符号整数,负数编码更小 |
| `sint64` | `int64` | zigzag varint | 同上,64 位 |
| `fixed32` | `uint32` | 4 bytes | 固定 4 字节,适合频繁序列化的场景 |
| `fixed64` | `uint64` | 8 bytes | 固定 8 字节 |
| `sfixed32` | `int32` | 4 bytes | 有符号固定 4 字节 |
| `sfixed64` | `int64` | 8 bytes | 有符号固定 8 字节 |
| `bool` | `bool` | varint (0/1) | — |
| `string` | `string` | len-delimited | UTF-8 编码 |
| `bytes` | `[]byte` | len-delimited | 任意二进制数据 |
### 性能选型建议
下面展示两个典型场景:
```go
// ❌ 不推荐:盲目使用 int64 增加序列化体积
// 每个 int64 可能占用 10+ bytes(varint 随数值增长)
message Request {
int64 user_id = 1; // 2^31 ≈ 21 亿,99% 的用户 ID 不会超过
int64 amount = 2; // float 存金额会丢失精度,且编码更大
}
// ✅ 推荐:按实际范围选型
message Request {
int32 user_id = 1; // 大多数用户 ID < 2^31
int32 amount_cents = 2; // 以"分"为单位存,避免 float,更节省
}
// ⭐ 极端优化:正负波动且范围小的场景
message Offset {
sint32 delta = 1; // zigzag 编码,-1 只占 1 byte(int32 需 5 byte)
}
```
上面的代码对应三种策略:
1. **默认选择 `int32`**:覆盖 ±21 亿的范围,对于 ID、计数等绝大多数场景足够。
2. **金额用最小货币单位存为整数**:比如 `100` 代表 ¥1.00,避免 IEEE 754 精度损失。
3. **`sint32` 用于小范围正负波动**:如 offset、delta,zigzag 编码让 `-1` 和 `1` 都只需 1 byte。
> [!tip] float vs double 取舍
> HTTP/2 + TLS 已经压缩了网络传输,**节省几个字节对延迟的影响微乎其微**。优先选择 `float32`,除非你的业务需要 IEEE 754 双精度精度(如金融计算)。
### Varint 编码与 zigzag 的关系
很多人分不清 varint 和 zigzag,这里简单拆解:
```mermaid
graph LR
A["原始整数"] --> B{"是否为负数?"}
B -- 否 --> C["varint: 每 7 bits 一组, MSB 标记 continuation"]
B -- 是 --> D["zigzag: n → (n << 1) ^ (n >> 31)"]
D --> C
C --> E["变长字节序列: 小数字仅 1 byte"]
```
- **varint**:只处理非负数,数字越小占的字节越少。`1` 占 1 byte,`2^31` 占 5 bytes。
- **zigzag**:将有符号整数映射为非负数,公式 `(n << 1) ^ (n >> 31)`,让 `-1` 变成 `1`,`-2` 变成 `3`,从而也能用 varint 紧凑编码。
## Repeated 与 Packed
`repeated` 字段表示一个动态长度的列表。在 proto3 中,numeric 类型的 repeated 默认采用 **packed encoding**(打包编码),非 numeric 类型(如 string、message)只能是 unpacked:
```protobuf
message TagList {
repeated string tags = 1; // string 类型无法 packed,总是 len-delimited
repeated int32 scores = 2; // int32 默认 packed
repeated float32 weights = 3; // float 默认 packed(4-byte fixed)
}
```
### Packed Encoding 原理与对比
考虑一组 `repeated int32` 字段 `[1, 2, 3]`,两种编码方式的 wire format 对比:
```mermaid
block
column "Unpacked (legacy)"
B1["tag(1B)"] B2["val 1(1B)"] B3["tag(1B)"] B4["val 2(1B)"] B5["tag(1B)"] B6["val 3(1B)"]
style B1 fill:#f9d
style B3 fill:#f9d
style B5 fill:#f9d
note1["重复写 tag\n共 6 bytes"]
column "Packed (proto3 默认)"
C1["tag(1B)"] C2["len(1B)"] C3["val 1(1B)"] C4["val 2(1B)"] C5["val 3(1B)"]
style C1 fill:#9df
style C2 fill:#9df
style C3 fill:#dfd
style C4 fill:#dfd
style C5 fill:#dfd
note2["只写一次 tag\n共 5 bytes"]
```
随着元素数量增长,差距越来越明显:
| 元素数量 | Unpacked | Packed | 节省比例 |
|---------|----------|--------|---------|
| 3 | 6B | 5B | 17% |
| 10 | 20B | 12B | 40% |
| 100 | 200B | 109B | 46% |
| 1000 | 2000B | 1037B | 48% |
> [!note] 手动关闭 packed
> 如果出于兼容性考虑需要关闭 packed,可以在 proto2 中使用:
> ```protobuf
> repeated int32 scores = 1 [packed = false]; // proto2 语法
> ```
> proto3 不允许此属性(必须 packed)。
## Wrapper Types 包装类型
Proto3 移除了 `required` 后,引入了 `google.protobuf.*_wrapper` 类型来区分「未设置」和「零值」:
```protobuf
import "google/protobuf/wrappers.proto";
message UserUpdate {
string id = 1;
google.protobuf.StringValue display_name = 2; // 可选的字符串
google.protobuf.BoolValue is_active = 3; // 可选的布尔值
google.protobuf.Int32Value age = 4; // 可选的整数
google.protobuf.FloatValue height_cm = 5; // 可选的浮点数
}
```
在 Go 生成的代码中,wrapper 类型生成的是**指针**:
```go
type UserUpdate struct {
Id string
DisplayName *string // nil = 未设置;"" = 明确设为空串
IsActive *bool // nil = 未设置;*false = 明确设为 false
Age *int32 // nil = 未设置;0 = 明确设为 0
}
```
这样就能清晰表达三种状态:**没传这个字段**(nil)、**传了但值是零**(指向零值的指针)、**传了正常值**(指向非零值的指针)。
### 原始类型 vs Wrapper 类型对比
| 场景 | 原始类型 `string` | Wrapper `StringValue` |
|------|-------------------|----------------------|
| Go 零值 | `""`(与"未设置"无法区分) | `nil`(清晰表达缺失) |
| JSON 序列化 | `"name": ""` | `"name": null` 或省略 |
| 判断是否传值 | 需要额外逻辑 | `if v != nil` 即可 |
| wire 大小 | 同左 | 同左(额外一层 wrapper overhead ≈ 0) |
> [!warning] Wrapper 不是银弹
> 不要把所有字段都用 wrapper。只有在 **你需要区分"未设置"和"零值"** 时才用 wrapper,否则会增加 nil-check 的心智负担。
### 哪些 Wrapper 可用
Protobuf 提供了所有标量类型的 wrapper,Go 中一一对应:
| Wrapper Type | Go 指针类型 | 典型用途 |
|-------------|-----------|---------|
| `StringValue` | `*string` | 可选文本 |
| `BoolValue` | `*bool` | 可选开关 |
| `Int32Value` | `*int32` | 可选小整数 |
| `Int64Value` | `*int64` | 可选大整数 / 时间戳 |
| `FloatValue` | `*float32` | 可选浮点 |
| `DoubleValue` | `*float64` | 可选双精度 |
| `BytesValue` | `*[]byte` | 可选二进制数据 |
## Well-Known Types
Protobuf 内置了一组通用的消息类型,称为 Well-Known Types(WKT),全部定义在 `google/protobuf/` 下。它们在不同语言中有各自的 native 映射,是实现跨语言兼容的关键。
核心 WWT 分类如下:
```mermaid
graph LR
A["Well-Known Types"] --> B["日期/时间\nTimestamp / Duration"]
A --> C["可选包装\nWrapper Types × 7"]
A --> D["泛型/动态\nAny / Value / Struct"]
A --> E["实用工具\nFieldMask / Empty / ..."]
```
### 时间相关:Timestamp & Duration
```protobuf
import (
"google/protobuf/timestamp.proto"
"google/protobuf/duration.proto"
)
message Task {
string title = 1;
google.protobuf.Timestamp deadline = 2; // 绝对时间点
google.protobuf.Duration timeout = 3; // 相对时长
}
```
在 Go 端,这两个类型直接映射为 `time.Time` 和 `time.Duration`,无需手动转换:
```go
task := &pb.Task{
Title: "发布版本",
Deadline: timestamppb.Now(), // 自动转当前 time.Time
Timeout: durationpb.New(30*time.Second), // 自动转 30s
}
```
> [!important] Timestamp 的序列化差异
> 在 JSON 映射中,`Timestamp` 默认序列化为 `RFC3339` 格式的 string:`"2026-05-11T08:30:00Z"`。但在 binary protobuf 中,它是两个 int64:seconds + nanoseconds。**跨语言调用时需确保对方也理解这种语义**。
### FieldMask:精准 Partial Update
`FieldMask` 是 gRPC 生态中最被低估的 WKT 之一。配合 `google.golang.org/protobuf/proto` 提供的 `ApplyFieldMask` 函数,可以实现精准的增量更新:
```protobuf
import "google/protobuf/field_mask.proto";
message UserPatchRequest {
google.protobuf.FieldMask update_mask = 1; // ["display_name", "email"]
User user = 2;
}
```
```go
// 服务器端:只对 mask 中指定的字段做更新
updatedUser := &existingUser
proto.ApplyFieldMask(&updatedUser, req.GetUser())
```
JSON 传递时也很简洁:`{ "updateMask": "display_name,email", "user": { "display_name": "新名字" } }`。
### Any:泛型消息容器
`Any` 允许你在不知道具体消息类型的情况下传递消息,常用于事件总线或插件架构:
```protobuf
import "google/protobuf/any.proto";
message Event {
google.protobuf.Any payload = 1; // 任意 protobuf message
}
```
反序列化时需要注册 type registry:
```go
// 注册已知类型
ptypes.RegisterAnyType(reflect.TypeFor[OrderCreated]())
// 从 Any 中提取具体类型
event := &Event{}
payload, _ := ptypes.UnmarshalAny(event.Payload)
```
> [!danger] 谨慎使用 Any
> `Any` 绕过了静态类型检查,滥用会导致调试困难。只在**真正的扩展点**(如插件系统、事件溯源)使用,不要用它来替代正常的消息设计。
## Map 类型细节
Map 在 wire format 中被编码为 `repeated key_value message`,底层实现其实就是一个 repeated:
```protobuf
message UserPreferences {
map<string, string> theme_settings = 1;
map<int32, string> role_permissions = 2;
}
```
关键行为:
- **迭代顺序不保证**:JSON/binary 序列化后顺序不可预测,不能依赖顺序做比较。
- **不能有嵌套 map**:`map<string, map<string,int>>` 非法。
- **key 只能是整数或字符串**,不支持 message 类型作为 key。
- Go 中初始值为 `nil`(而非 `make(map[string]string)`),使用前需判空或初始化。
### Map vs Message + repeated
当需要额外元数据时,map 就不够用了,需要改用 message + repeated:
```protobuf
// ❌ map 只能存 key-value,无法携带额外信息
message Bad {
map<string, string> roles = 1;
}
// ✅ 用 message 承载完整信息
message Good {
message RoleMapping {
string role = 1;
string permission = 2;
}
repeated RoleMapping mappings = 1;
}
```
## Optional 与 Oneof
回到开头的问题:proto3 没有 `required`,如何保证字段一定存在?
**方案一:Wrapper Type**(见上文)— 适合"可选但可零值"的场景。
**方案二:Oneof** — 适合"多个字段中必须有且仅有一个"的场景:
```protobuf
message PaymentRequest {
string order_id = 1;
oneof payment_method {
string alipay_token = 2;
string wechat_pay_nonce = 3;
string bank_card_number = 4;
}
}
```
在 Go 生成的代码中,oneof 会生成一个接口来标识哪个字段被设置了:
```go
// Go 端生成的 interface
type PaymentRequest_PaymentMethod interface {
isPaymentRequest_PaymentMethod()
}
```
使用时通过类型断言判断:
```go
switch req.GetPaymentMethod().(type) {
case *PaymentRequest_AlipayToken:
// 走支付宝
case *PaymentRequest_WechatPayNonce:
// 走微信支付
default:
// 错误:payment method 未设置
}
```
> [!example] Oneof 的实际应用场景
> - **多态请求参数**:搜索时可以按关键词、ID 或模糊匹配,三者选一
> - **协议切换**:同一个连接支持多种子协议
> - **互斥配置**:比如渲染模式只能选一种(WebGL / Canvas / SVG)
## 最佳实践总结
- **优先使用 `int32`**,除非确定数据范围超过 ±21 亿才用 `int64`。
- **金额相关用整型存储最小货币单位**(如 cents),永远不要用 `float` 存钱。
- **需要表达"可选"时优先考虑 wrapper types**,比 oneof 更简洁,比裸 scalar 更能区分零值和缺失。
- **timestamp 统一用 RFC3339 string**,跨语言互通性最好。
- **sint32/sint64** 仅在小范围内有正负波动的场景(如 offset、delta)中使用。
- **oneof 用在"多选一"的互斥场景**,而不是用来模拟 optional。
- **FieldMask 是实现 RESTful PATCH 语义的神器**,别自己解析 JSON 路径了。
- **慎用 Any**,只在真正的扩展点使用,避免绕过类型安全。
## 关联笔记
- [[hhs/gRPC/1. Protobuf 基础篇/01-Protobuf 语法与消息定义]] — Protobuf 语法入门,建议先读本篇再来看本文
- [[hhs/gRPC/1. Protobuf 基础篇/03-字段编号与前向兼容]] — 字段编号管理、向前向后兼容规则
- [[hhs/gRPC/1. Protobuf 基础篇/04-Oneof 与包装类型]] — Oneof 深度使用 + Wrapper Type 实战模式