Reading

Papers & Books

Technical literature I've read, am reading, or plan to. Mostly distributed systems, databases, data engineering, and AI. Notes are my own.

Distributed Systems

read ★★★★★

The paper that defined large-scale batch processing. The simplicity of the programming model is the insight.

read ★★★★★

Foundational for understanding wide-column stores and the architecture of HBase, Cassandra.

Dynamo: Amazon's Highly Available Key-Value Store DeCandia et al. — Amazon · 2007
read ★★★★★

Eventual consistency, vector clocks, consistent hashing. Still the best paper on availability-first design.

Spanner: Google's Globally Distributed Database Corbett et al. — Google · 2012
read ★★★★★

TrueTime and externally-consistent distributed transactions. Remarkable engineering.

read ★★★★★

Much easier to understand than Paxos. Good foundation for understanding etcd, CockroachDB.

Paxos Made Simple Lamport · 2001
read ★★★★☆

Still not that simple, but essential reading. Read Raft first.

Designing Data-Intensive Applications Martin Kleppmann · 2017
read ★★★★★

The best systems book written in the last decade. Required reading for anyone building data systems.

Databases

WiredTiger: A High Performance, Multi-Version Concurrency Control Storage Engine WiredTiger Team — MongoDB · 2014
read ★★★★☆

Understand MVCC and B-tree vs LSM-tree tradeoffs. Explains much of MongoDB's storage behavior.

read ★★★★☆

Good read on how a production storage engine evolves under real workloads.

Apache Pinot: Realtime OLAP for 530 Million Users LinkedIn Engineering · 2018
read ★★★★☆

Architecture of how Pinot handles real-time + offline data for low-latency analytics at scale.

read ★★★★★

Best resource for understanding database indexing across engines. Practical and deep.

Data Engineering

read ★★★★★

The original Kafka paper. The log abstraction is the key insight. Still holds up.

The conceptual foundation for Flink, Beam, and modern stream processing. Windowing, triggers, accumulation.

ACID on top of Parquet in object storage. Foundational for understanding modern lakehouse architectures.

Spark: Cluster Computing with Working Sets Zaharia et al. · 2010
read ★★★★☆

The original Spark paper. RDD model and in-memory computation for iterative algorithms.

AI & Machine Learning

Attention Is All You Need Vaswani et al. — Google Brain · 2017
read ★★★★★

The transformer architecture paper. Everything in modern AI traces back here.

DeepSeek-V3 Technical Report DeepSeek AI · 2024
read ★★★★★

MLA, DeepSeekMoE, FP8 training. Remarkable efficiency. Changed the conversation on inference cost.

read ★★★★☆

Good detail on training data curation, instruction tuning, and safety filtering at scale.

Mixtral of Experts Mistral AI · 2024
read ★★★★☆

Sparse mixture-of-experts applied to an open model. MoE becomes practical.

read ★★★★☆

The CoT paper. Step-by-step reasoning significantly improves complex task performance.