F · English–Chinese Glossary

Every term from the small term tables (术语对照) at the end of the Prologue, the 24 chapters and the Epilogue, in one A–Z list, plus a few terms that only the appendices use. Search in English or in Chinese, or jump to a letter. The Taught in link opens the place that first teaches the term: a chapter’s term table, or an appendix (marked App. A, App. B or App. E). When chapters give different Chinese words for one term, each word is shown with its chapter (why).

A

English 中文 Taught in Also in
A/A test A/A 测试 Ch 21 –
A/B test A/B 测试 Prologue Ch 20
absence of evidence 缺乏证据(不等于没有) Ch 18 –
absolute difference 绝对差值 Ch 20 –
accuracy 准确性 Ch 15 –
ACID 事务特性(原子性、一致性、隔离性、持久性) Ch 11 –
active customer 活跃客户 Ch 5 Ch 6
adaptive query execution 自适应查询执行 Ch 10 –
additive / semi-additive / non-additive fact 可加 / 半可加 / 不可加度量 Ch 12 –
ADS (application data service) 应用层 Ch 12 –
aggregate 聚合 Ch 3 –
aging schedule 账龄分析表 Ch 6 –
alias 别名 Ch 3 –
allowed lateness 允许延迟 Ch 14 –
alpha spending α 消耗 Ch 22 –
alternative hypothesis 备择假设 Ch 18 –
always-valid p-value 始终有效的 p 值 Ch 22 –
analytical procedures 分析程序 Ch 4 Ch 16
annotation 标注 Ch 5 –
anomaly detection 异常检测 Ch 15 –
answer first / bottom line up front (BLUF) 结论先行 Ch 24 –
anti-join 反连接 Ch 3 Ch 7
app user (opened the app) 应用用户 Ch 5 –
append-only log 仅追加日志 Ch 8 –
append / overwrite 追加写 / 覆盖写 Ch 13 –
appendix 附录 Ch 24 –
arm 实验分组(臂) Ch 20 –
as-of time 数据截止时点 Ch 1 –
as originally reported 原报告数 Ch 11 –
assignment / exposure 分组 / 曝光 Ch 20 –
at-least-once 至少一次 Ch 8 –
at-most-once 至多一次 Ch 8 –
atomic metric / derived metric 原子指标 / 派生指标 Ch 12 –
average treatment effect (ATE) 平均处理效应 App. B –
axis 坐标轴 Ch 5 –

B

English 中文 Taught in Also in
B+ tree B+树 Ch 9 –
B-tree B树(也写作 B-树;连字符不是“减号”) Ch 9 –
back-test 回测 Ch 15 –
back to the table 回表 Ch 9 –
backfill 补数 / 回填 Ch 13 –
background loss 背景丢失(正常丢失率) Ch 15 –
bandwidth (regression discontinuity) 带宽 App. B –
baseline 基线 / 对比基准 (Prologue)
基线 (Ch 19)
Prologue Ch 19
batch control 批次控制 Ch 13 –
batch processing 批处理 Ch 14 –
Bayesian 贝叶斯 Ch 22 –
before / after image 前镜像 / 后镜像 Ch 7 –
bell shape / normal distribution 钟形 / 正态分布 Ch 16 –
Benjamini–Hochberg procedure BH 校正(Benjamini–Hochberg 方法) App. B –
binlog (binary log) 二进制日志 Ch 7 –
bitmap 位图 App. E –
blameless review 无责复盘 Ch 24 –
block 数据块 Ch 10 –
block bootstrap 区组自助法 / 块自助法 Ch 17 –
Bonferroni correction Bonferroni 校正 App. B –
bootstrap 自助法 Ch 17 –
bounded / unbounded data 有界 / 无界数据 Ch 14 –
break-even 盈亏平衡点 Ch 24 –
breakdown point 崩溃点 Ch 2 –
bridge chart / waterfall chart 桥图 / 瀑布图 Ch 24 –
broadcast join (map join) 广播连接(map join) Ch 10 –
bucketing 分桶 App. E –
business (operational) database 业务库 Ch 7 –

C

English 中文 Taught in Also in
CAATs 计算机辅助审计技术 Ch 3 –
catalog 目录 Ch 11 –
causal diagram 因果图 Ch 23 –
causal inference 因果推断 Ch 23 –
central limit theorem 中心极限定理 App. E –
change data capture (CDC) 变更数据捕获 Ch 7 –
change table / change log 变更表 / 变更日志 Ch 12 –
chargeback 拒付 / 退单 Ch 11 –
checkpoint 检查点 Ch 14 –
chi-square test 卡方检验 App. B –
chimney development (each team builds its own tables) 烟囱式开发 App. E –
churn 客户流失 Ch 6 –
client-side / server-side 客户端 / 服务端 Ch 7 –
cluster / node 集群 / 节点 Ch 10 –
clustered index 聚簇索引 Ch 9 –
cohort 同期群 Ch 6 –
cohort table / heatmap 同期群表 / 热力图 Ch 6 –
column store 列存 Ch 9 –
combiner 合并器(map 端预聚合) Ch 10 –
commit 提交 Ch 11 –
committed offset 已提交偏移量 Ch 8 –
compaction 合并 / 压实 Ch 11 –
compaction (LSM tree) 合并 / 压缩合并 Ch 9 –
compaction (small files) 小文件合并 Ch 9 –
comparison group 对照组 / 比较组 Ch 23 –
completeness 完整性 Prologue Ch 15
completeness / occurrence 完整性 / 发生性 Ch 7 –
composite index 联合索引 Ch 9 –
compression 压缩 Ch 9 –
confidence interval 置信区间 Ch 17 –
confidence level 置信水平 Ch 17 –
conformed dimension 一致性维度 Ch 12 –
confounder 混杂因素 Ch 4 Ch 16
confounding 混杂 Ch 23 –
consistency 一致性 Ch 15 –
consumer 消费者 Ch 8 –
consumer group 消费者组 Ch 8 –
control 控制 Ch 15 –
control total 控制总数 Ch 13 –
control / treatment 对照组 / 实验组 Ch 20 –
controlling for 控制(某变量) Ch 23 –
conversion 转化率 Ch 19 –
conversion rate 转化率 Ch 4 –
copy-on-write 写时复制 Ch 11 –
corporate account 企业客户 Ch 2 –
correlation 相关系数 Ch 22 –
corroborating evidence 佐证 Ch 23 –
counter-metric 反向指标 / 制衡指标 App. E –
counterfactual 反事实 Ch 23 –
covariate 协变量 Ch 22 –
covering index 覆盖索引 Ch 9 –
credible interval 可信区间 Ch 22 –
criteria, condition, root cause, effect 标准、现状、根本原因、影响(审计发现要素) Ch 24 –
CTE (common table expression) 公用表表达式 Ch 3 –
CUPED CUPED(利用实验前数据的方差缩减) Ch 22 –
cut-off 截止 / 截止性测试 Ch 14 –
cut-off testing 截止性测试 Ch 11 –
cutoff 断点 / 阈值 Ch 23 –

D

English 中文 Taught in Also in
DAG (directed acyclic graph) 有向无环图 Ch 13 –
daily full copy of a table 每日全量快照 Ch 12 –
dashboard 仪表盘 / 看板 Prologue –
data contract 数据契约 Ch 15 –
data lake 数据湖 Ch 11 –
data locality 数据本地性 Ch 10 –
data pipeline 数据管道 Ch 13 –
data quality 数据质量 Ch 15 –
data skew 数据倾斜 Ch 10 –
data swamp 数据沼泽 Ch 11 –
data test 数据测试 Ch 15 –
data warehouse 数据仓库 Prologue Ch 11
decision rule 决策规则 Ch 20 –
definition 口径 Ch 1 –
degrees of freedom 自由度 App. B –
deletion vector 删除向量 Ch 11 –
delta method Delta 方法 Ch 21 –
dependency 依赖 Ch 13 –
design / operating effectiveness 设计有效性 / 执行有效性 Ch 15 –
dictionary encoding 字典编码 Ch 9 –
difference-in-differences (DiD) 双重差分 Ch 23 –
DIM (dimension layer) 维度层 Ch 12 –
dimension table 维度表 Ch 12 –
direct label 直接标注 Ch 5 –
disaggregation 分解 / 细分 Ch 4 –
distributed 分布式 Ch 10 –
distribution 分布 Ch 2 –
driver / executor 驱动器 / 执行器 Ch 10 –
drop-off 步骤流失 / 掉队 Ch 6 –
dual axis (dual-axis chart) 双坐标轴(双轴图) Ch 5 –
dump 全量导出 Ch 7 –
DWD (data warehouse detail) 明细层 Ch 12 –
DWS (data warehouse summary) 汇总层 Ch 12 –

E

English 中文 Taught in Also in
effect size 效应量 Ch 19 –
ETL / ELT 抽取-转换-加载 / 抽取-加载-转换 Ch 13 –
event 事件 Ch 1 Ch 7
event / message 事件 / 消息 Ch 8 –
event time 事件时间 Ch 14 –
event time / record time 事件时间 / 记录时间 Ch 7 –
exactly-once 精确一次 Ch 8 –
exaggeration ratio 夸大比 Ch 19 –
executive summary 执行摘要 Ch 24 –
expectation (expected value) 预期值(统计学中称期望值) Ch 16 –
expected loss 期望损失 Ch 22 –
external table 外部表 Ch 10 –

F

English 中文 Taught in Also in
fact table 事实表 Ch 12 –
fair presentation 公允列报 Ch 5 –
false discovery rate (FDR) 错误发现率 Ch 21 –
false positive 假阳性 Ch 21 –
family-wise error rate (FWER) 族错误率 Ch 21 –
fan-out (B+ tree) 扇出 Ch 9 –
fan-out (join) 扇出 / 重复计数 Ch 3 –
freshness 新鲜度 / 时效性 Ch 13 –
funnel 漏斗 Ch 6 –

G

English 中文 Taught in Also in
gaps and islands (consecutive days) 连续登录 / 间隙与岛屿 App. A –
grain 粒度 Ch 1 Ch 3, Ch 12
gross merchandise value (GMV) 商品交易总额 / 成交总额 Ch 12 –
GROUP BY 分组 Ch 3 –
guardrail metric 护栏指标 Ch 20 Ch 21

H

English 中文 Taught in Also in
hash index 哈希索引 Ch 9 –
hash total 哈希总数 Ch 13 –
HDFS 分布式文件系统 Ch 10 –
headline 标题句 Ch 24 –
high watermark (Kafka) 高水位 App. E –
histogram 直方图 Ch 2 –
holdout 保留组 Ch 20 –
hot key 热点键 Ch 10 –
hot partition 热点分区 Ch 8 –
HyperLogLog HyperLogLog(基数估计算法) App. E –
hypothesis 假设 Ch 18 Ch 20

I

English 中文 Taught in Also in
idempotent 幂等 Ch 8 Ch 13
in-sync replicas (ISR) 同步副本集合 App. E –
independent 独立 Ch 17 –
index 索引 Ch 9 –
index (base = 100) 指数(基期 = 100) Ch 5 –
index condition pushdown 索引下推 Ch 9 –
information fraction 信息比例 Ch 22 –
ingest time 接收时间 Ch 7 –
inner join / left join 内连接 / 左连接 Ch 3 –
interference 干扰 / 溢出效应 Ch 21 –
IT general controls 信息技术一般控制 Ch 9 –

J

English 中文 Taught in Also in
join 连接(表关联) (Ch 1)
连接 (Ch 3)
Ch 1 Ch 3
journey memo 流水账式报告 Ch 24 –
judgement 判断 Ch 24 –

K

English 中文 Taught in Also in
key (message key) 键(消息键) Ch 8 –
key (primary key) 主键 Ch 1 –
KPI (key performance indicator) 关键绩效指标 Ch 2 –
KPI pack 关键绩效指标报告 Ch 5 –

L

English 中文 Taught in Also in
label 标签 Ch 1 –
lag 消费积压 / 延迟 Ch 8 –
lakehouse 湖仓一体 Ch 11 –
Lambda / Kappa architecture Lambda / Kappa 架构 Ch 14 –
late data 迟到数据 Ch 14 –
latency 延迟 Ch 14 –
law of large numbers 大数定律 Ch 16 –
lazy evaluation 惰性求值 Ch 10 –
least squares 最小二乘法 Ch 16 –
leftmost prefix 最左前缀 Ch 9 –
legend 图例 Ch 5 –
lie factor 失真因子 Ch 5 –
lineage 血缘 Ch 12 Ch 15
linear regression 线性回归 Ch 16 –
logical date / data date 逻辑日期 / 业务日期 Ch 13 –
long tail 长尾 Ch 2 –
LSM tree LSM树 Ch 9 –

M

English 中文 Taught in Also in
managed (internal) table 管理表(内部表) Ch 10 –
Mann–Whitney U test 曼-惠特尼 U 检验 App. E –
MapReduce MapReduce(映射-归约) Ch 10 –
margin 毛利(每杯利润) Ch 24 –
margin of error 误差范围 Ch 17 –
matching 匹配 Ch 23 –
materiality 重要性 Ch 19 –
mean 均值 Ch 2 –
measure 度量 Ch 12 –
measurement 度量 / 观测值 Ch 1 –
medallion architecture (bronze / silver / gold) 奖章架构(铜 / 银 / 金) Ch 12 –
median 中位数 Ch 2 –
merge-on-read 读时合并 Ch 11 –
message queue 消息队列 Ch 8 –
metadata 元数据 Ch 11 –
metastore 元数据存储 Ch 10 –
metric 指标 Prologue Ch 1
metric dictionary 指标字典 Ch 12 –
metric layer / semantic layer 指标层 / 语义层 Ch 12 –
metric tree 指标树 App. E –
minimum detectable effect (MDE) 最小可检测效应 Ch 19 –
miss (residual) 残差 Ch 18 –
mix 结构 / 构成 Ch 4 –
mix shift / composition effect 结构变化 / 构成效应 Ch 4 –
mode 众数 Ch 2 –
modifier / time period 修饰词 / 时间周期 Ch 12 –
MPP (massively parallel processing) 大规模并行处理 Ch 11 –
mSPRT (mixture sequential probability ratio test) 混合序贯概率比检验 App. B –
multiple comparisons 多重比较 Ch 21 –

N

English 中文 Taught in Also in
NameNode / DataNode 名称节点 / 数据节点 Ch 10 –
narrow / wide dependency (Spark) 窄依赖 / 宽依赖 App. E –
natural experiment 自然实验 Ch 23 –
negative assurance 消极保证 Ch 18 –
noise 噪声 Ch 16 –
non-sampling risk 非抽样风险 Ch 17 –
north-star metric 北极星指标 App. E –
novelty effect 新奇效应 Ch 20 Ch 21
NULL 空值 Ch 3 –
null distribution 零分布 Ch 18 –
null hypothesis 原假设 / 零假设 Ch 18 –

O

English 中文 Taught in Also in
O’Brien–Fleming / Pocock boundary O’Brien–Fleming / Pocock 边界 Ch 22 –
object storage 对象存储 Ch 10 –
observational data 观测数据 Ch 23 –
ODS (operational data store) 贴源层 / 操作数据层 Ch 12 –
offset 偏移量 Ch 8 –
OLAP (online analytical processing) 联机分析处理 Ch 9 –
OLTP (online transaction processing) 联机事务处理 Ch 9 –
one-sided / two-sided test 单侧 / 双侧检验 Ch 18 –
outage 故障 / 宕机 Ch 24 –
outlier 异常值 Ch 2 –
overall conversion rate 整体转化率 Ch 6 –
overlap 共同支撑 / 重叠 Ch 23 –
owner 负责人 Ch 15 –

P

English 中文 Taught in Also in
p-value p 值 Ch 18 –
page 页 Ch 9 –
parallel trends 平行趋势 Ch 23 –
partition 分区 Ch 8 Ch 9, Ch 13
partition pruning 分区裁剪 Ch 9 –
peeking 偷看 / 提前看结果 Ch 21 –
percentage point 百分点 Prologue Ch 1, Ch 4
percentile 分位数 / 百分位数 (Ch 2)
百分位数 (Ch 17)
Ch 2 Ch 17
percentile interval (bootstrap) 百分位数区间 App. B –
periodic snapshot fact table 周期快照事实表 Ch 12 –
pickup code 取餐码 Prologue –
pivot / unpivot 行转列 / 列转行 App. E –
placebo test 安慰剂检验 Ch 23 –
planted secret 预埋的线索 Epilogue –
population 总体 Ch 1 Ch 16
posterior 后验 Ch 22 –
power 统计功效 Ch 18 Ch 19
practical significance 实际显著性 Ch 19 –
pre-mortem 事前验尸 Ch 24 –
pre-period 实验前时期 Ch 22 –
pre-registration 预注册 Ch 20 Ch 21
precision (share of flagged items that are real) 精确率 / 查准率 App. E –
predicate pushdown 谓词下推 Ch 9 –
primary metric 主指标 Ch 20 –
prior 先验 Ch 22 –
processing time 处理时间 Ch 14 –
producer 生产者 Ch 8 –
production 生产环境 Ch 9 –
propensity score 倾向得分 Ch 23 –
property 属性 Ch 7 –
pyramid principle 金字塔原理 Ch 24 –

Q

English 中文 Taught in Also in
query 查询 Ch 1 Ch 3
query engine 查询引擎 Ch 11 –

R

English 中文 Taught in Also in
ramp-up 逐步放量 Ch 20 –
random sample 随机样本 Ch 16 –
random stream 随机数流 Epilogue –
randomisation unit 随机化单元 Ch 20 –
range 区间 Ch 24 –
ratio metric 比率指标 Ch 21 –
read replica 只读副本 Ch 9 –
real-time data warehouse 实时数仓 App. E –
rebalance (consumer group) 重平衡 / 再均衡 App. E –
recommendation 建议 Ch 24 –
reconcile, reconciliation 对账 / 核对 Ch 1 –
reconciliation 对账 Ch 7 Ch 15
record count 记录数 Ch 13 –
regression discontinuity (RD) 断点回归 Ch 23 –
regression to the mean 均值回归 Ch 16 –
relative lift 相对提升 Ch 20 –
replay 重放 (Ch 7)
重放 / 回溯 (Ch 8)
Ch 7 Ch 8
replication factor 副本因子 Ch 8 Ch 10
resample 重抽样 Ch 17 –
restatement 重述 Ch 11 –
retention 留存 Ch 6 –
retention curve 留存曲线 Ch 6 –
retention (log retention) 保留期 Ch 8 –
retry 重试 Ch 13 –
reversible decision 可逆决策 Ch 24 –
row group / column chunk 行组 / 列块 Ch 9 –
row store 行存 Ch 9 –
run-length encoding 游程编码 Ch 9 –
run log 运行日志 Ch 13 –
running total 累计值 Ch 3 –
running variable 驱动变量 Ch 23 –

S

English 中文 Taught in Also in
salting 加盐 Ch 10 –
sample 样本 Ch 16 –
sample ratio mismatch (SRM) 样本比例失衡 Ch 20 Ch 21
sample size 样本量 Ch 19 –
sampling risk 抽样风险 Ch 17 –
savepoint 保存点 Ch 14 –
scale up / scale out 纵向扩展 / 横向扩展 Ch 10 –
scheduler 调度器 Ch 13 –
schema evolution 模式演进 Ch 11 –
schema-on-read 读时模式 Ch 10 –
seasonality 季节性 Ch 16 –
secondary index 二级索引 Ch 9 –
secondary metric 次要指标 Ch 20 –
seed 随机种子 Epilogue –
segment 细分群体 (Ch 2)
细分 / 分群 (Ch 4)
Ch 2 Ch 4
segmentation 细分分析 Ch 4 –
selection 选择偏差 Ch 23 –
semi join 半连接 App. A –
sequential testing 序贯检验 Ch 22 –
serial / parallel complement 串行补数 / 并行补数 Ch 13 –
server log 服务端日志 Ch 7 –
session 会话 Ch 6 –
severity 严重级别 Ch 15 –
sharp / fuzzy regression discontinuity 精确断点回归 / 模糊断点回归 App. B –
shuffle 洗牌 / 数据重分布 Ch 10 –
side output 侧输出 Ch 14 –
significance level (alpha) 显著性水平 Ch 18 –
Simpson’s paradox 辛普森悖论 Ch 4 –
skew 偏态 Ch 2 –
SLA (service level agreement) 服务等级协议 Ch 13 –
slowly changing dimension (SCD) 缓慢变化维 Ch 12 –
small files 小文件 Ch 9 –
small multiples 小多图 Ch 5 –
snapshot 快照 Ch 11 –
snowflake schema 雪花模型 Ch 12 –
source 数据源 Ch 7 –
source document 原始凭证 Ch 7 –
stakeholder 利益相关方 Ch 24 –
standard deviation 标准差 Ch 16 –
standard deviation (SD) 标准差 Ch 17 –
standard error (SE) 标准误 Ch 17 –
standardisation 标准化(按相同结构比较) Ch 4 –
STAR method (situation, task, action, result) STAR 法则 App. E –
star schema 星型模型 Ch 12 –
state 状态 Ch 14 –
statistical significance 统计显著性 Ch 19 –
statistically significant 统计显著 Ch 18 –
statistics (min/max) 统计信息(最小值/最大值) Ch 9 –
status 状态 Ch 1 –
step conversion rate 步骤转化率 Ch 6 –
stratified sampling 分层抽样 Ch 2 –
stream processing 流处理 Ch 14 –
surrogate key 代理键 Ch 12 –
switchback test 轮转实验 Ch 21 –
synthetic data 合成数据 Epilogue –
system of record / book of record 记录系统 / 权威数据源 Ch 7 –

T

English 中文 Taught in Also in
t-test t 检验 App. E –
table format 表格式 Ch 11 –
table / row / column 表 / 行 / 列 Ch 3 –
tails 尾部 Ch 2 –
task 任务 Ch 13 –
task / stage 任务 / 阶段 Ch 10 –
test (pytest) 测试 Epilogue –
test statistic 检验统计量 Ch 18 –
threshold 阈值 Ch 15 –
time travel 时间旅行 Ch 11 –
time zone 时区 Ch 1 –
timeliness 及时性 Ch 15 –
tolerable misstatement 可容忍错报 Ch 17 Ch 19
topic 主题 Ch 8 –
traceability 可追溯性 Ch 24 –
tracing 顺查(从凭证到账) Prologue –
tracking / tracking plan 埋点 / 埋点方案 Ch 7 –
trend 趋势 Ch 16 –
trimmed mean 截尾均值 Ch 2 –
truncated axis 截断坐标轴 Ch 5 –
tumbling / sliding / session window 滚动 / 滑动 / 会话窗口 Ch 14 –
two-phase commit 两阶段提交 Ch 14 –
Type I error 第一类错误 Ch 18 –
Type II error 第二类错误 Ch 18 –

U

English 中文 Taught in Also in
uncertainty 不确定性 Ch 24 –
underpowered 功效不足 Ch 19 –
uniqueness 唯一性 Ch 15 –
unit 计量单位 Ch 1 –
upsert 更新插入 Ch 13 –
upsert / merge 更新插入 / 合并 Ch 11 –
UTC 协调世界时 Ch 1 –

V

English 中文 Taught in Also in
validity 有效性 Ch 15 –
variable cost 变动成本 Ch 24 –
variance 方差 Ch 16 –
variance reduction 方差缩减 Ch 22 –
vouching 逆查(从账到凭证) Prologue –

W

English 中文 Taught in Also in
walkthrough 穿行测试 Ch 6 –
watermark 水位线 Ch 14 –
week over week 周环比 Prologue –
weighted average 加权平均 Ch 4 –
Welch’s t-test Welch t 检验 App. B –
wide table 宽表 Ch 12 –
Wilson interval 威尔逊区间 App. B –
window 窗口 Ch 14 –
window function 窗口函数 Ch 3 –
winner’s curse 赢家诅咒 Ch 19 –
winsorize 缩尾 Ch 21 –
with replacement 有放回 Ch 17 –
working papers 审计工作底稿 Ch 24 –
write amplification 写放大 Ch 9 –
write-up 实验结论报告 Ch 20 –

Z

English 中文 Taught in Also in
zipper table 拉链表 Ch 12 –

About this list

The list holds 451 terms. 419 come from the 452 rows of the 26 chapter term tables; 30 of these appear in more than one chapter. The other 32 are terms that the appendices use and no chapter table has; they come from a short list kept by hand.

When two chapters give different Chinese words for one English term, the table shows each word with its chapter. Five terms are like this. Usually both words are in common use: percentile, for example, is 分位数 or 百分位数. When one English word has two meanings, the chapters name them apart, as with key (primary key) and key (message key).

The list is built from the chapters each time the book is built, so it always matches them.