G · Further Reading
Every book, paper and documentation page that the chapters recommend, in one place. If you want one next book, start with the short shelf below. The full list follows, grouped by the Part of this book where each reference first appears; under each one you can see which chapters cite it. A DOI (digital object identifier) is a permanent ID for a paper or book; its link opens the publisher’s page.
Start here
Seven books for three paths. Each one is the natural next step after a part of this book.
For data analysis
- The Art of Statistics, David Spiegelhalter (Pelican, 2019). How to learn from data through real questions, with very few formulas. Chapter 24 cites it; read it alongside Part III. Cited in Ch 24.
- Fundamentals of Data Visualization, Claus O. Wilke (O’Reilly Media, 2019; free to read online). How to choose a chart for your data and draw it clearly and honestly. It goes further than Chapter 5. Cited in Ch 5.
- Storytelling with Data, Cole Nussbaumer Knaflic (Wiley, 2015). How to give one chart one clear message for a busy reader. It pairs well with the memo in Chapter 24. Cited in Ch 24.
For data engineering
- Designing Data-Intensive Applications, Martin Kleppmann (O’Reilly Media, 2017). How data systems store data, copy it, split it across machines and process it in batches and streams, and what each design gives up. It explains in depth how the systems in Part II work. Chinese edition: 《数据密集型应用系统设计》. Not cited in a chapter.
- The Data Warehouse Toolkit, Ralph Kimball and Margy Ross (3rd edition, Wiley, 2013). The standard book on facts, dimensions, grain and slowly changing dimensions: the long version of Chapter 12. Cited in Ch 1, Ch 12.
For experimentation
- Trustworthy Online Controlled Experiments, Ron Kohavi, Diane Tang and Ya Xu (Cambridge University Press, 2020). The standard practical guide to A/B testing in a company: units, metrics, trust checks, and much more than Chapters 19–22 can cover. Cited in Ch 19, Ch 20, Ch 21, Ch 22.
- Causal Inference: The Mixtape, Scott Cunningham (Yale University Press, 2021; the author also hosts a free online version). Difference-in-differences, regression discontinuity, matching and more, with worked code. The next step after Chapter 23. Cited in Ch 23.
One paper to add: Deng, Xu, Kohavi and Walker (2013), the CUPED paper (DOI). Cited in Ch 22.
Everything the chapters cite
The chapters’ Further reading lists hold 94 entries. Some works are cited in more than one chapter, so the list below has 89 different references (3 of them are cited more than once), plus one core book that no chapter cites. Each reference is listed under the Part where it first appears.
Part I · Look Before You Leap
- Deng, A., & Shi, X. (2016). Data-Driven Metric Development for Online Controlled Experiments: Seven Lessons Learned. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 77–86. DOI. How a large company designs, checks and changes its metric definitions. Cited in Ch 1.
- Dmitriev, P., Gupta, S., Kim, D. W., & Vaz, G. (2017). A Dirty Dozen: Twelve Common Metric Interpretation Pitfalls in Online Controlled Experiments. Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1427–1436. DOI. Twelve common ways to misread a metric’s movement, with real cases from a large company. Cited in Ch 1.
- Ralph Kimball and Margy Ross, The Data Warehouse Toolkit: The Definitive Guide to Dimensional Modeling, 3rd edition, Wiley, 2013. The standard book on facts, dimensions, grain and slowly changing dimensions. Cited in Ch 1, Ch 12.
- Anscombe, F. J. (1973). Graphs in Statistical Analysis. The American Statistician, 27(1), 17–21. DOI. Four small data sets with the same summary numbers and very different shapes: the classic argument for drawing your data. Cited in Ch 2.
- Matejka, J., & Fitzmaurice, G. (2017). Same Stats, Different Graphs: Generating Datasets with Varied Appearance and Identical Statistics through Simulated Annealing. Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems, 1290–1294. DOI. A modern, playful version of Anscombe’s point. Cited in Ch 2.
- von Hippel, P. T. (2005). Mean, Median, and Skew: Correcting a Textbook Rule. Journal of Statistics Education, 13(2). DOI. Why “mean above median” is not a safe test for skew. Cited in Ch 2.
- E. F. Codd, “A relational model of data for large shared data banks”, Communications of the ACM 13(6), 377–387, 1970. doi:10.1145/362384.362685. The short, famous paper behind the idea of tables, rows and keys. Cited in Ch 3.
- The DuckDB documentation on the SELECT statement and on window functions. DuckDB is the database inside this book and inside its playgrounds. Cited in Ch 3.
- Simpson, E. H. (1951). The Interpretation of Interaction in Contingency Tables. Journal of the Royal Statistical Society: Series B, 13(2), 238–241. doi:10.1111/j.2517-6161.1951.tb00088.x. The short paper the paradox is named after. Cited in Ch 4.
- Bickel, P. J., Hammel, E. A., & O’Connell, J. W. (1975). Sex Bias in Graduate Admissions: Data from Berkeley. Science, 187(4175), 398–404. doi:10.1126/science.187.4175.398. The most famous real case: a university’s total admission rates seemed to favour men, but department by department there was no such bias. Women had applied more often to the departments that admitted a smaller share of applicants. Cited in Ch 4.
- Pearl, J. (2014). Comment: Understanding Simpson’s Paradox. The American Statistician, 68(1), 8–13. doi:10.1080/00031305.2014.876829. Why the right answer depends on what causes what, not on the numbers alone. Cited in Ch 4.
- Edward R. Tufte, The Visual Display of Quantitative Information (Graphics Press, 1983; second edition 2001). The classic book on honest statistical graphics, and the source of the lie factor. Cited in Ch 5.
- William S. Cleveland and Robert McGill, “Graphical Perception: Theory, Experimentation, and Application to the Development of Graphical Methods”, Journal of the American Statistical Association 79(387), 531–554, 1984. doi:10.1080/01621459.1984.10478080. The experiments behind “position beats angle”. Cited in Ch 5.
- Claus O. Wilke, Fundamentals of Data Visualization (O’Reilly Media), free to read online. A practical guide to choosing and drawing charts. Cited in Ch 5.
- Fader, P. S., & Hardie, B. G. S. (2007). How to Project Customer Retention. Journal of Interactive Marketing, 21(1), 76–90. doi:10.1002/dir.20074. A short, practical paper with a spreadsheet model, written for subscription businesses, where you can see each customer cancel. It shows why a group’s retention tends to level off: customers differ, and the ones most likely to cancel cancel first. Cited in Ch 6.
Part II · The Journey of One Order
- The MySQL Reference Manual, “The Binary Log” and “Replication Formats”: what the binlog holds, and statement-based versus row-based logging. Cited in Ch 7.
- The Debezium documentation, “Debezium connector for MySQL”: how a CDC tool reads the binlog, and what its change events look like. Cited in Ch 7.
- Jay Kreps, “The Log: What every software engineer should know about real-time data’s unifying abstraction”, LinkedIn Engineering blog, 2013. A long, friendly essay by one of Kafka’s creators on why the log sits at the centre of data systems. Cited in Ch 8.
- The Apache Kafka documentation, especially the Introduction and the Design section (“Message Delivery Semantics” and “Replication”). Cited in Ch 8.
- The MySQL 8.4 Reference Manual on indexes: Clustered and Secondary Indexes, How MySQL Uses Indexes and Multiple-Column Indexes. Cited in Ch 9.
- Rudolf Bayer and Edward McCreight, “Organization and maintenance of large ordered indexes”, Acta Informatica 1(3), 1972. doi:10.1007/BF00288683. The paper that introduced the B-tree. Cited in Ch 9.
- Douglas Comer, “The ubiquitous B-tree”, ACM Computing Surveys 11(2), 1979. doi:10.1145/356770.356776. A readable survey, including the B+ tree. Cited in Ch 9.
- Patrick O’Neil, Edward Cheng, Dieter Gawlick and Elizabeth O’Neil, “The log-structured merge-tree (LSM-tree)”, Acta Informatica 33(4), 1996. doi:10.1007/s002360050048. Cited in Ch 9.
- The Apache Parquet documentation, especially Concepts and the File Format pages on metadata, encodings and the page index. Cited in Ch 9.
- The Apache ORC documentation, especially Indexes and the ORC v1 specification. Cited in Ch 9.
- Daniel J. Abadi, Samuel R. Madden and Nabil Hachem, “Column-stores vs. row-stores: how different are they really?”, Proceedings of the 2008 ACM SIGMOD International Conference on Management of Data. doi:10.1145/1376616.1376712. A careful study of why column stores are fast. Cited in Ch 9.
- Jeffrey Dean and Sanjay Ghemawat, “MapReduce: simplified data processing on large clusters”, Communications of the ACM 51(1), 2008. doi:10.1145/1327452.1327492. Cited in Ch 10.
- Ashish Thusoo and others, “Hive: a warehousing solution over a map-reduce framework”, Proceedings of the VLDB Endowment 2(2), 2009. doi:10.14778/1687553.1687609. Cited in Ch 10.
- Matei Zaharia and others, “Apache Spark: a unified engine for big data processing”, Communications of the ACM 59(11), 2016. doi:10.1145/2934664. Cited in Ch 10.
- The official documentation: the HDFS Architecture guide, the Spark RDD Programming Guide (section “Shuffle operations”), and Spark’s Performance Tuning page (adaptive query execution). Cited in Ch 10.
- Michael Armbrust and others, “Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores”, Proceedings of the VLDB Endowment 13(12), 2020, pages 3411–3424. doi:10.14778/3415478.3415560. How a log of commits turns files in object storage into ACID tables. Cited in Ch 11.
- Michael Armbrust, Ali Ghodsi, Reynold Xin and Matei Zaharia, “Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics”, CIDR 2021. The paper that set out the lakehouse idea. Cited in Ch 11.
- The Apache Iceberg table specification: snapshots, manifests and commits, written for implementers but readable. Cited in Ch 11.
- The Apache Hudi documentation, “Table Types”: copy-on-write and merge-on-read, with diagrams. Cited in Ch 11.
- Kimball Group, “Dimensional Modeling Techniques”. Short official pages on the four-step process, grain, star schemas, snowflaking and each SCD type. Cited in Ch 12.
- Alibaba Cloud DataWorks documentation: “Data warehouse layering” (ODS, DIM, DWD, DWS and ADS) and “Data metric” (atomic metrics, modifiers, time periods and derived metrics). Cited in Ch 12.
- Databricks documentation, “What is the medallion lakehouse architecture?”. Bronze, silver and gold, as Databricks describes them. Cited in Ch 12.
- Apache Airflow documentation: “Best Practices” (tasks as transactions, no plain
INSERT, read and write fixed partitions), “DAG Runs” (data intervals, logical dates, catchup and backfill) and “Tasks” (retries, task states, and the move from SLAs to deadline alerts). Cited in Ch 13. - Apache DolphinScheduler documentation, “Workflow Definition”: running a workflow with serial or parallel complement. Cited in Ch 13.
- Alibaba Cloud DataWorks documentation, “Get started with Operation Center”: data backfill and baseline monitoring. Cited in Ch 13.
- Apache Hive documentation, “LanguageManual DML”:
INSERT OVERWRITEandINSERT INTO, with partitions. Cited in Ch 13. - Maxime Beauchemin, “Functional Data Engineering: a modern paradigm for batch data processing”, 2018. An essay by the creator of Airflow on pure, idempotent tasks and partitions that are overwritten, never edited. Cited in Ch 13.
- The Apache Flink documentation: “Timely Stream Processing” (event time and watermarks), “Windows” (window types, allowed lateness, side outputs), and “Stateful Stream Processing” (state and checkpoints). Cited in Ch 14.
- Tyler Akidau and others, “The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing”, Proceedings of the VLDB Endowment 8(12), 2015, pages 1792–1803. doi:10.14778/2824032.2824076. The paper behind many of the ideas in Chapter 14. Cited in Ch 14.
- Piotr Nowojski and Mike Winters, “An Overview of End-to-End Exactly-Once Processing in Apache Flink (with Apache Kafka, too!)”, Apache Flink blog, 2018. The two-phase commit, step by step. Cited in Ch 14.
- Jay Kreps, “Questioning the Lambda Architecture”, O’Reilly Radar, 2014. The essay that named the Kappa architecture. Cited in Ch 14.
- Richard Y. Wang and Diane M. Strong, “Beyond Accuracy: What Data Quality Means to Data Consumers”, Journal of Management Information Systems 12(4), 1996, pages 5–33. doi:10.1080/07421222.1996.11518099. The classic study that asked data users which qualities matter to them. Cited in Ch 15.
- Leo L. Pipino, Yang W. Lee and Richard Y. Wang, “Data Quality Assessment”, Communications of the ACM 45(4), 2002, pages 211–218. doi:10.1145/505248.506010. Short and practical: how to turn quality dimensions into measurements. Cited in Ch 15.
- The dbt documentation, “Add data tests to your DAG”: data tests as queries that return failing rows, and the four built-in generic tests. Cited in Ch 15.
- Kleppmann, M. (2017). Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems. O’Reilly Media. Book website. How data systems store, copy, split up and process data, and what each design gives up. Chinese edition: 《数据密集型应用系统设计》 (中国电力出版社, 2018). Not cited in a chapter: a core book for all of Part II.
Part III · Living With Uncertainty
- Galton, F. (1886). Regression Towards Mediocrity in Hereditary Stature. The Journal of the Anthropological Institute of Great Britain and Ireland, 15, 246–263. doi:10.2307/2841583. Where the name “regression” comes from. Cited in Ch 16.
- Barnett, A. G., van der Pols, J. C., & Dobson, A. J. (2005). Regression to the mean: what it is and how to deal with it. International Journal of Epidemiology, 34(1), 215–220. doi:10.1093/ije/dyh299. A short, clear guide with examples. Cited in Ch 16.
- Kahneman, D., & Tversky, A. (1973). On the psychology of prediction. Psychological Review, 80(4), 237–251. doi:10.1037/h0034747. Includes the famous flight instructors who believed that praise makes pilots worse. Cited in Ch 16.
- International Auditing and Assurance Standards Board. ISA 520, Analytical Procedures. IAASB Handbook, 2012 edition (PDF). Paragraphs 5 and 7 are the four steps and the duty to investigate. Cited in Ch 16.
- Efron, B. (1979). Bootstrap Methods: Another Look at the Jackknife. The Annals of Statistics, 7(1), 1–26. doi:10.1214/aos/1176344552. The paper that introduced the bootstrap. Cited in Ch 17.
- Künsch, H. R. (1989). The Jackknife and the Bootstrap for General Stationary Observations. The Annals of Statistics, 17(3), 1217–1241. doi:10.1214/aos/1176347265. The block bootstrap with overlapping blocks, for data that are linked over time. Cited in Ch 17.
- Carlstein, E. (1986). The Use of Subseries Values for Estimating the Variance of a General Statistic from a Stationary Sequence. The Annals of Statistics, 14(3). doi:10.1214/aos/1176350057. Non-overlapping blocks, like Mia’s weeks. Cited in Ch 17.
- Neyman, J. (1937). Outline of a Theory of Statistical Estimation Based on the Classical Theory of Probability. Philosophical Transactions of the Royal Society of London. Series A, 236(767), 333–380. doi:10.1098/rsta.1937.0005. Where confidence intervals come from. Cited in Ch 17.
- Cumming, G., & Finch, S. (2005). Inference by Eye: Confidence Intervals and How to Read Pictures of Data. American Psychologist, 60(2), 170–180. doi:10.1037/0003-066X.60.2.170. Includes the rule for reading two overlapping intervals. Cited in Ch 17.
- Hoekstra, R., Morey, R. D., Rouder, J. N., & Wagenmakers, E.-J. (2014). Robust misinterpretation of confidence intervals. Psychonomic Bulletin & Review, 21(5), 1157–1164. doi:10.3758/s13423-013-0572-3. Researchers and students misread intervals in the same ways; a useful warning. Cited in Ch 17.
- International Auditing and Assurance Standards Board. ISA 530, Audit Sampling. IAASB Handbook, 2012 edition (PDF). Paragraph 5 defines sampling risk and non-sampling risk. Cited in Ch 17.
- Wasserstein, R. L., & Lazar, N. A. (2016). The ASA Statement on p-Values: Context, Process, and Purpose. The American Statistician, 70(2), 129–133. DOI Cited in Ch 18.
- Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. European Journal of Epidemiology, 31(4), 337–350. DOI Cited in Ch 18.
- Phipson, B., & Smyth, G. K. (2010). Permutation P-values Should Never Be Zero: Calculating Exact P-values When Permutations Are Randomly Drawn. Statistical Applications in Genetics and Molecular Biology, 9(1). DOI Cited in Ch 18.
- Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press. DOI. The standard practical guide: the unit, metrics and guardrails, trust checks, ramps and holdouts. Cited in Ch 19, Ch 20, Ch 21, Ch 22.
- Gelman, A., & Carlin, J. (2014). Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors. Perspectives on Psychological Science, 9(6), 641–651. DOI Cited in Ch 19.
- Button, K. S., Ioannidis, J. P. A., Mokrysz, C., Nosek, B. A., Flint, J., Robinson, E. S. J., & Munafò, M. R. (2013). Power failure: why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 14(5), 365–376. DOI Cited in Ch 19.
- Kohavi, R., Deng, A., & Vermeer, L. (2022). A/B Testing Intuition Busters. Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 3168–3177. DOI Cited in Ch 19.
Part IV · The Art of the Experiment
- Johari, R., Koomen, P., Pekelis, L., & Walsh, D. (2017). Peeking at A/B Tests: Why It Matters, and What to Do About It. Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1517–1525. DOI. Always-valid p-values and the mSPRT. Cited in Ch 21, Ch 22.
- Fabijan, A., Gupchup, J., Gupta, S., Omhover, J., Qin, W., Vermeer, L., & Dmitriev, P. (2019). Diagnosing Sample Ratio Mismatch in Online Controlled Experiments. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. DOI Cited in Ch 21.
- Holm, S. (1979). A Simple Sequentially Rejective Multiple Test Procedure. Scandinavian Journal of Statistics, 6, 65–70. Cited in Ch 21.
- Benjamini, Y., & Hochberg, Y. (1995). Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B, 57(1), 289–300. DOI Cited in Ch 21.
- Deng, A., Knoblich, U., & Lu, J. (2018). Applying the Delta Method in Metric Analytics. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. DOI Cited in Ch 21.
- Bojinov, I., Simchi-Levi, D., & Zhao, J. (2023). Design and Analysis of Switchback Experiments. Management Science, 69(7). DOI Cited in Ch 21.
- Deng, A., Xu, Y., Kohavi, R., & Walker, T. (2013). Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data. Proceedings of the Sixth ACM International Conference on Web Search and Data Mining (WSDM), 123–132. DOI. The CUPED paper. Cited in Ch 22.
- Card, D., & Krueger, A. B. (1994). Minimum Wages and Employment: A Case Study of the Fast-Food Industry in New Jersey and Pennsylvania. American Economic Review, 84(4), 772–793. RePEc Cited in Ch 23.
- Angrist, J. D., & Pischke, J.-S. (2009). Mostly Harmless Econometrics: An Empiricist’s Companion. Princeton University Press. DOI Cited in Ch 23.
- Imbens, G. W., & Lemieux, T. (2008). Regression Discontinuity Designs: A Guide to Practice. Journal of Econometrics, 142(2), 615–635. DOI Cited in Ch 23.
- Bertrand, M., Duflo, E., & Mullainathan, S. (2004). How Much Should We Trust Differences-in-Differences Estimates? The Quarterly Journal of Economics, 119(1), 249–275. DOI Cited in Ch 23.
- Conley, T. G., & Taber, C. R. (2011). Inference with “Difference in Differences” with a Small Number of Policy Changes. The Review of Economics and Statistics, 93(1), 113–125. DOI Cited in Ch 23.
- Rosenbaum, P. R., & Rubin, D. B. (1983). The Central Role of the Propensity Score in Observational Studies for Causal Effects. Biometrika, 70(1), 41–55. DOI Cited in Ch 23.
- Cunningham, S. (2021). Causal Inference: The Mixtape. Yale University Press. DOI. The author’s free online version is at mixtape.scunning.com, which now hosts the second edition, Causal Inference: The Remix, still in progress. Cited in Ch 23.
Part V · From Numbers to Decisions
- Minto, B. (1996). The Minto Pyramid Principle: Logic in Writing, Thinking and Problem Solving. Minto Books International. Publisher. The source of “answer first, then the reasons, grouped under it”; the publisher notes translations into Chinese and other languages. Cited in Ch 24.
- van der Bles, A. M., van der Linden, S., Freeman, A. L. J., Mitchell, J., Galvao, A. B., Zaval, L., & Spiegelhalter, D. J. (2019). Communicating uncertainty about facts, numbers and science. Royal Society Open Science, 6(5), 181870. DOI. A careful review of how to say how sure you are, and what happens when you do. Cited in Ch 24.
- Spiegelhalter, D. (2020). The Art of Statistics: Learning from Data (paperback; first published 2019). Pelican. Publisher. A plain-language tour of learning from data, by a statistician who writes for the public. Cited in Ch 24.
- Kent, S. Words of Estimative Probability. Studies in Intelligence, 8(4). Central Intelligence Agency, Center for the Study of Intelligence. CIA. The essay behind the “serious possibility” story: what words like “probable” mean to different readers. Cited in Ch 24.
- The Institute of Internal Auditors (2024). Global Internal Audit Standards. IIA. Standards 14.2–14.4 and 15.1 describe how a finding is built and reported. Cited in Ch 24.
- Klein, G. (2007). Performing a Project Premortem. Harvard Business Review, September 2007. HBR. The pre-mortem in two pages. Cited in Ch 24.
- Knaflic, C. N. (2015). Storytelling with Data: A Data Visualization Guide for Business Professionals. Wiley. DOI. How to choose one chart and make its point easy to see. Cited in Ch 24.
- Lunney, J., & Lueder, S. (2016). Postmortem Culture: Learning from Failure. In B. Beyer, C. Jones, J. Petoff, & N. R. Murphy (Eds.), Site Reliability Engineering. O’Reilly Media. Free online. How a large engineering team writes reviews that look for causes, not culprits. Cited in Ch 24.