Guided reading tracks for specific goals, a topic index, and the prerequisite map for every chapter — so you can chart a path instead of reading front to back.
Guided reading tracks
ML engineers building a pretraining stackTrain an LLM From Scratch
For engineers who want to go from raw text to a trained, sampling checkpoint. You'll build every piece of the pipeline: tokenizer, transformer, pretraining objective, scaling plan, and the end-to-end training recipe.
- 1.6 Neural Networks From Scratch: MLPs & Backprop
- 1.7 Automatic Differentiation & PyTorch Internals
- 2.1 Tokenization: BPE, WordPiece, Unigram & Byte-Level
- 2.3 The Attention Mechanism From Scratch
- 2.4 Multi-Head Attention, MQA, GQA & MLA
- 2.5 Positional Encodings: Sinusoidal, Learned, RoPE & ALiBi
- 2.6 The Transformer Block: Norms, Residuals, MLPs & Activations
- 2.7 Building a GPT From Scratch (nanoGPT-style)
- 3.1 Pretraining Data: Sources, Crawling & The Data Pipeline
- 3.3 The Pretraining Objective & Loss
- 3.4 Scaling Laws: Kaplan, Chinchilla & Beyond
- 3.9 Optimizers: SGD, Adam, Adafactor, Lion, Muon & Shampoo
- 3.17 Training an LLM From Scratch: The End-to-End Recipe
- 14.1 The Capstone: Building Stack-100M, and the 2026 Small-Model Landscape
Inference/serving engineersInference & Serving Optimization
For engineers serving LLMs in production who need low latency and high throughput. You'll master inference anatomy, continuous batching, KV-cache management, quantization, and the modern serving stacks.
- 2.4 Multi-Head Attention, MQA, GQA & MLA
- 7.1 The Anatomy of LLM Inference: Prefill, Decode & The KV Cache
- 7.2 Continuous Batching & Request Scheduling
- 4.6 PagedAttention & KV-Cache Memory Management
- 7.3 vLLM: Architecture, PagedAttention & Internals
- 7.7 Prefix Caching & KV-Cache Reuse
- 7.4 SGLang: RadixAttention & Structured Programs
- 7.6 Speculative Decoding: Draft Models, Medusa, EAGLE & Lookahead
- 4.7 Quantization I: Post-Training Quantization (GPTQ, AWQ, SmoothQuant)
- 4.8 Quantization II: INT4/INT8/FP8, GGUF, bitsandbytes & QAT
- 7.9 Sampling Strategies & Decoding Algorithms
- 7.5 TensorRT-LLM, TGI & Other Serving Stacks
- 7.12 Inference Economics: Latency, Throughput & Cost
- 12.1 Designing an LLM Serving System
Alignment & post-training engineersPost-Training & Alignment
For engineers turning a base model into a helpful, aligned assistant. You'll work through SFT, reward modeling, PPO/DPO/GRPO, RL with verifiable rewards, the failure modes, and how to evaluate the result.
- 5.1 Supervised Fine-Tuning & Instruction Tuning
- 5.2 Chat Templates, Data Formatting & Sequence Packing
- 5.3 PEFT I: LoRA, QLoRA, DoRA & The Adapter Family
- 5.5 The RLHF Pipeline & Reward Modeling
- 5.6 Policy Gradients & PPO for Language Models
- 5.7 Direct Preference Optimization & Its Variants
- 5.8 GRPO, RLOO & Critic-Free RL
- 5.9 RL with Verifiable Rewards (RLVR) & The Reasoning Recipe
- 6.1 The Anatomy of an RL-for-LLM System
- 5.13 Reward Hacking, Over-Optimization & Alignment Failures
- 11.1 The Evaluation Problem & Benchmark Landscape
- 11.2 LLM-as-a-Judge & Automated Evaluation
Agent & applied-LLM engineersAgents & Tool Use
For engineers building LLM agents that call tools, retrieve knowledge, and coordinate. You'll cover tool use, agentic loops, harness and context engineering, memory, RAG, and multi-agent orchestration.
- 8.1 Tool Use & Function Calling
- 8.2 The Agentic Loop: ReAct, Plan-Execute & Reflection
- 8.3 Harness Engineering: Building a Coding Agent
- 8.4 Context Engineering & Management
- 8.5 Memory Systems for Agents
- 8.6 The Model Context Protocol (MCP)
- 9.1 Embeddings & Representation Learning
- 9.2 Vector Databases & Approximate Nearest Neighbor Search
- 9.3 Retrieval-Augmented Generation Architectures
- 9.4 Chunking, Reranking & Hybrid Search
- 9.5 Advanced RAG: GraphRAG, Agentic RAG & Long-Context vs RAG
- 8.7 Multi-Agent Systems & Orchestration
- 8.8 Agent Evaluation & Benchmarks
GPU/performance & systems engineersGPU Systems & Performance
For performance engineers who want to make LLMs go fast on real hardware. You'll cover GPU architecture, the roofline model, CUDA and Triton kernels, FlashAttention, compilers, and distributed training.
- 1.8 GPU Architecture & The Memory Hierarchy
- 4.1 The Roofline Model & Performance Engineering
- 4.5 CUDA Programming Essentials for ML Engineers
- 4.4 Writing GPU Kernels with Triton
- 4.2 FlashAttention I: IO-Awareness & The Online Softmax
- 4.3 FlashAttention 2 & 3: Work Partitioning, Warp Specialization & FP8
- 4.9 Kernel Fusion, torch.compile, CUDA Graphs & Compilers
- 4.10 Memory-Efficient Training: Checkpointing, Offloading & LoRA Math
- 1.9 Parallel Computing & Collective Communication
- 3.5 Distributed Training I: Data Parallelism, DDP, ZeRO & FSDP
- 3.6 Distributed Training II: Tensor, Pipeline, Sequence & Expert Parallelism
- 3.7 Megatron-LM, DeepSpeed & Parallelism in Practice
ML/LLM engineer interview candidatesThe ML Interview Fast Path
For candidates prepping for an ML/LLM engineer interview. A high-yield sweep of the concepts most likely to come up: fundamentals, the transformer, pretraining and scaling, alignment, inference, kernels, RAG, agents, and evaluation.
- 1.5 Machine Learning Fundamentals
- 1.2 Probability, Statistics & Information Theory
- 2.3 The Attention Mechanism From Scratch
- 2.4 Multi-Head Attention, MQA, GQA & MLA
- 2.6 The Transformer Block: Norms, Residuals, MLPs & Activations
- 2.7 Building a GPT From Scratch (nanoGPT-style)
- 3.3 The Pretraining Objective & Loss
- 3.4 Scaling Laws: Kaplan, Chinchilla & Beyond
- 5.1 Supervised Fine-Tuning & Instruction Tuning
- 5.7 Direct Preference Optimization & Its Variants
- 7.1 The Anatomy of LLM Inference: Prefill, Decode & The KV Cache
- 4.2 FlashAttention I: IO-Awareness & The Online Softmax
- 9.3 Retrieval-Augmented Generation Architectures
- 8.2 The Agentic Loop: ReAct, Plan-Execute & Reflection
- 11.2 LLM-as-a-Judge & Automated Evaluation
Browse by topic
alignment (22)
1.2 Probability, Statistics & Information Theory 3.15 Synthetic Data for Pre- and Post-Training 5.1 Supervised Fine-Tuning & Instruction Tuning 5.5 The RLHF Pipeline & Reward Modeling 5.6 Policy Gradients & PPO for Language Models 5.7 Direct Preference Optimization & Its Variants 5.8 GRPO, RLOO & Critic-Free RL 5.9 RL with Verifiable Rewards (RLVR) & The Reasoning Recipe 5.11 Constitutional AI, RLAIF & Self-Improvement 5.13 Reward Hacking, Over-Optimization & Alignment Failures 6.3 TRL: HuggingFace's RL Library 6.5 OpenRLHF, NeMo-Aligner & Ray-Based Systems 6.9 Advantage Estimation, KL Control & Stability Tricks 11.2 LLM-as-a-Judge & Automated Evaluation 11.5 Red-Teaming, Safety & Robustness Evaluation 12.5 Data Flywheels & Continuous Improvement 13.2 Knowledge Editing & Machine Unlearning 13.5 AI Safety: Scalable Oversight, Dangerous-Capability Evals & Frontier Safety 14.9 Post-Training: SFT, DPO, and Narrow RLVR (GRPO) That Works at 100M Glossary of Terms Key Papers: An Annotated Reading List From-Scratch Code Index
architecture (36)
1.1 Linear Algebra for Deep Learning 1.6 Neural Networks From Scratch: MLPs & Backprop 1.8 GPU Architecture & The Memory Hierarchy 1.10 The Accelerator Landscape: TPUs, Trainium, AMD/ROCm & Gaudi 2.1 Tokenization: BPE, WordPiece, Unigram & Byte-Level 2.2 Embeddings & The Input Pipeline 2.3 The Attention Mechanism From Scratch 2.4 Multi-Head Attention, MQA, GQA & MLA 2.5 Positional Encodings: Sinusoidal, Learned, RoPE & ALiBi 2.6 The Transformer Block: Norms, Residuals, MLPs & Activations 2.7 Building a GPT From Scratch (nanoGPT-style) 2.8 Architecture Variants: Encoder-Decoder, Decoder-Only & Prefix-LM 2.9 Mixture-of-Experts (MoE) Architectures 2.10 Modern Architecture Improvements & Design Choices 2.11 Beyond Attention: SSMs, Mamba, RWKV & Linear Attention 2.12 Diffusion & Non-Autoregressive Language Models 3.6 Distributed Training II: Tensor, Pipeline, Sequence & Expert Parallelism 3.13 Long-Context Pretraining & Context Extension 3.16 Continual & Domain-Adaptive Pretraining 4.5 CUDA Programming Essentials for ML Engineers 5.4 PEFT II: Prompt/Prefix Tuning, IA3, Model Merging & Soups 9.1 Embeddings & Representation Learning 10.1 Vision Transformers & Image Encoders 10.2 Vision-Language Models 10.3 Audio, Speech & Multimodal Fusion 10.4 Diffusion Models & Generative Modeling (Breadth) 10.5 Unified & Any-to-Any Models 13.1 Mechanistic Interpretability & Model Internals 14.1 The Capstone: Building Stack-100M, and the 2026 Small-Model Landscape 14.3 A Byte-Level BPE Tokenizer From Scratch (and Why Vocab Size Is a Design Lever at 100M) 14.4 The Stack-100M Architecture: SOTA Components, Cited and Assembled 14.8 Mid-Training: Quality Annealing, Long-Context Extension & Capability Injection Glossary of Terms The Math Reference Sheet Key Papers: An Annotated Reading List From-Scratch Code Index
data (24)
1.5 Machine Learning Fundamentals 2.1 Tokenization: BPE, WordPiece, Unigram & Byte-Level 3.1 Pretraining Data: Sources, Crawling & The Data Pipeline 3.2 Data Cleaning, Deduplication & Quality Filtering 3.3 The Pretraining Objective & Loss 3.4 Scaling Laws: Kaplan, Chinchilla & Beyond 3.14 Data Mixing, Domain Weighting & Curriculum 3.15 Synthetic Data for Pre- and Post-Training 3.16 Continual & Domain-Adaptive Pretraining 5.1 Supervised Fine-Tuning & Instruction Tuning 5.2 Chat Templates, Data Formatting & Sequence Packing 6.12 RL Data, Curriculum & Replay Management 9.2 Vector Databases & Approximate Nearest Neighbor Search 9.4 Chunking, Reranking & Hybrid Search 9.5 Advanced RAG: GraphRAG, Agentic RAG & Long-Context vs RAG 10.3 Audio, Speech & Multimodal Fusion 11.1 The Evaluation Problem & Benchmark Landscape 12.5 Data Flywheels & Continuous Improvement 13.3 Privacy, Memorization & Differential Privacy for LLMs 13.6 AI Governance, Compliance & Regulation 14.2 Data: Sourcing, Filtering, Dedup, Tokenize & Pack ~20B Tokens 14.3 A Byte-Level BPE Tokenizer From Scratch (and Why Vocab Size Is a Design Lever at 100M) 14.8 Mid-Training: Quality Annealing, Long-Context Extension & Capability Injection 14.10 A Narrow Auto-Research Agent: ReAct, Tool-Use & Retrieval by Distillation
evaluation (34)
1.2 Probability, Statistics & Information Theory 1.5 Machine Learning Fundamentals 3.2 Data Cleaning, Deduplication & Quality Filtering 3.15 Synthetic Data for Pre- and Post-Training 3.17 Training an LLM From Scratch: The End-to-End Recipe 5.5 The RLHF Pipeline & Reward Modeling 5.9 RL with Verifiable Rewards (RLVR) & The Reasoning Recipe 5.10 Reasoning, Chain-of-Thought & Test-Time Compute 5.13 Reward Hacking, Over-Optimization & Alignment Failures 6.8 Reward Engineering, Verifiers & Sandboxes 8.2 The Agentic Loop: ReAct, Plan-Execute & Reflection 8.3 Harness Engineering: Building a Coding Agent 8.8 Agent Evaluation & Benchmarks 8.9 Prompt Engineering as Engineering 9.1 Embeddings & Representation Learning 9.3 Retrieval-Augmented Generation Architectures 9.6 Multimodal & Visual-Document Retrieval: ColPali & Late Interaction 11.1 The Evaluation Problem & Benchmark Landscape 11.2 LLM-as-a-Judge & Automated Evaluation 11.3 Building Eval Harnesses 11.4 Reasoning, Coding & Agentic Evals 11.5 Red-Teaming, Safety & Robustness Evaluation 11.6 Statistical Rigor in Evaluation: Confidence Intervals & Significance 12.2 Observability, Logging & LLMOps 12.4 Safety, Guardrails & Content Moderation 12.5 Data Flywheels & Continuous Improvement 12.7 Online Evaluation: A/B Testing, Canaries & Guardrail Metrics 12.8 Reliability Engineering for LLM Systems: SLOs & Incident Response 13.1 Mechanistic Interpretability & Model Internals 13.3 Privacy, Memorization & Differential Privacy for LLMs 13.4 Watermarking, Provenance & AI-Content Detection 13.5 AI Safety: Scalable Oversight, Dangerous-Capability Evals & Frontier Safety 13.6 AI Governance, Compliance & Regulation 14.11 Evaluation & Serving: Honest Benchmarks, int4 Quantization, and Running on a Laptop
gpu (32)
1.4 Numerical Computing, Floating Point & Precision 1.7 Automatic Differentiation & PyTorch Internals 1.8 GPU Architecture & The Memory Hierarchy 1.9 Parallel Computing & Collective Communication 1.10 The Accelerator Landscape: TPUs, Trainium, AMD/ROCm & Gaudi 3.5 Distributed Training I: Data Parallelism, DDP, ZeRO & FSDP 3.6 Distributed Training II: Tensor, Pipeline, Sequence & Expert Parallelism 3.7 Megatron-LM, DeepSpeed & Parallelism in Practice 3.8 Mixed Precision, bf16 & FP8 Training 4.1 The Roofline Model & Performance Engineering 4.2 FlashAttention I: IO-Awareness & The Online Softmax 4.3 FlashAttention 2 & 3: Work Partitioning, Warp Specialization & FP8 4.4 Writing GPU Kernels with Triton 4.5 CUDA Programming Essentials for ML Engineers 4.6 PagedAttention & KV-Cache Memory Management 4.9 Kernel Fusion, torch.compile, CUDA Graphs & Compilers 4.10 Memory-Efficient Training: Checkpointing, Offloading & LoRA Math 6.3 TRL: HuggingFace's RL Library 6.4 veRL: HybridFlow & The Single-Controller Architecture 6.5 OpenRLHF, NeMo-Aligner & Ray-Based Systems 6.7 Colocated vs Disaggregated RL & Weight Synchronization 7.1 The Anatomy of LLM Inference: Prefill, Decode & The KV Cache 7.3 vLLM: Architecture, PagedAttention & Internals 7.4 SGLang: RadixAttention & Structured Programs 7.5 TensorRT-LLM, TGI & Other Serving Stacks 7.8 Disaggregated Prefill/Decode & Chunked Prefill 7.11 Multi-GPU & Multi-Node Inference 7.12 Inference Economics: Latency, Throughput & Cost 7.13 Serving Mixture-of-Experts: Expert Parallelism & All-to-All Inference 14.7 The Pretraining Run: A Complete Single-GPU Training Loop 14.12 Retrospective: Cost Accounting, Reproducibility, and the Path to 1B Tooling & Environment Setup Cheatsheet
inference (59)
1.10 The Accelerator Landscape: TPUs, Trainium, AMD/ROCm & Gaudi 2.3 The Attention Mechanism From Scratch 2.4 Multi-Head Attention, MQA, GQA & MLA 2.5 Positional Encodings: Sinusoidal, Learned, RoPE & ALiBi 2.7 Building a GPT From Scratch (nanoGPT-style) 2.8 Architecture Variants: Encoder-Decoder, Decoder-Only & Prefix-LM 2.10 Modern Architecture Improvements & Design Choices 2.11 Beyond Attention: SSMs, Mamba, RWKV & Linear Attention 2.12 Diffusion & Non-Autoregressive Language Models 3.13 Long-Context Pretraining & Context Extension 4.1 The Roofline Model & Performance Engineering 4.2 FlashAttention I: IO-Awareness & The Online Softmax 4.4 Writing GPU Kernels with Triton 4.6 PagedAttention & KV-Cache Memory Management 4.7 Quantization I: Post-Training Quantization (GPTQ, AWQ, SmoothQuant) 4.8 Quantization II: INT4/INT8/FP8, GGUF, bitsandbytes & QAT 4.9 Kernel Fusion, torch.compile, CUDA Graphs & Compilers 5.2 Chat Templates, Data Formatting & Sequence Packing 5.10 Reasoning, Chain-of-Thought & Test-Time Compute 5.12 Distillation, Model Compression & Knowledge Transfer 6.1 The Anatomy of an RL-for-LLM System 6.2 The Generation–Training Loop & Rollout Engines 6.7 Colocated vs Disaggregated RL & Weight Synchronization 6.10 Agentic & Multi-Turn RL 6.11 Scaling RL: Throughput, Load Balancing & The Latest Tricks 7.1 The Anatomy of LLM Inference: Prefill, Decode & The KV Cache 7.2 Continuous Batching & Request Scheduling 7.3 vLLM: Architecture, PagedAttention & Internals 7.4 SGLang: RadixAttention & Structured Programs 7.5 TensorRT-LLM, TGI & Other Serving Stacks 7.6 Speculative Decoding: Draft Models, Medusa, EAGLE & Lookahead 7.7 Prefix Caching & KV-Cache Reuse 7.8 Disaggregated Prefill/Decode & Chunked Prefill 7.9 Sampling Strategies & Decoding Algorithms 7.10 Structured & Constrained Generation 7.11 Multi-GPU & Multi-Node Inference 7.12 Inference Economics: Latency, Throughput & Cost 7.13 Serving Mixture-of-Experts: Expert Parallelism & All-to-All Inference 7.14 Multi-Tenant LoRA & Adapter Serving at Scale 8.1 Tool Use & Function Calling 8.2 The Agentic Loop: ReAct, Plan-Execute & Reflection 8.4 Context Engineering & Management 8.9 Prompt Engineering as Engineering 9.2 Vector Databases & Approximate Nearest Neighbor Search 9.6 Multimodal & Visual-Document Retrieval: ColPali & Late Interaction 10.2 Vision-Language Models 10.3 Audio, Speech & Multimodal Fusion 10.4 Diffusion Models & Generative Modeling (Breadth) 10.5 Unified & Any-to-Any Models 11.4 Reasoning, Coding & Agentic Evals 12.1 Designing an LLM Serving System 12.3 Caching, Routing & Cost Control in Production 14.4 The Stack-100M Architecture: SOTA Components, Cited and Assembled 14.5 Mini Scaling Laws: Fit Your Own Law Before Spending the Budget 14.11 Evaluation & Serving: Honest Benchmarks, int4 Quantization, and Running on a Laptop Glossary of Terms Key Papers: An Annotated Reading List Tooling & Environment Setup Cheatsheet From-Scratch Code Index
math (30)
1.1 Linear Algebra for Deep Learning 1.2 Probability, Statistics & Information Theory 1.3 Calculus, Optimization & Convexity 1.4 Numerical Computing, Floating Point & Precision 1.5 Machine Learning Fundamentals 1.6 Neural Networks From Scratch: MLPs & Backprop 1.7 Automatic Differentiation & PyTorch Internals 2.2 Embeddings & The Input Pipeline 2.3 The Attention Mechanism From Scratch 2.4 Multi-Head Attention, MQA, GQA & MLA 2.5 Positional Encodings: Sinusoidal, Learned, RoPE & ALiBi 2.6 The Transformer Block: Norms, Residuals, MLPs & Activations 2.11 Beyond Attention: SSMs, Mamba, RWKV & Linear Attention 3.2 Data Cleaning, Deduplication & Quality Filtering 3.3 The Pretraining Objective & Loss 3.4 Scaling Laws: Kaplan, Chinchilla & Beyond 3.9 Optimizers: SGD, Adam, Adafactor, Lion, Muon & Shampoo 4.2 FlashAttention I: IO-Awareness & The Online Softmax 5.6 Policy Gradients & PPO for Language Models 7.9 Sampling Strategies & Decoding Algorithms 8.8 Agent Evaluation & Benchmarks 9.1 Embeddings & Representation Learning 10.4 Diffusion Models & Generative Modeling (Breadth) 11.1 The Evaluation Problem & Benchmark Landscape 11.6 Statistical Rigor in Evaluation: Confidence Intervals & Significance 13.2 Knowledge Editing & Machine Unlearning 14.5 Mini Scaling Laws: Fit Your Own Law Before Spending the Budget 14.6 Optimizer & Schedule: Muon + MuonClip and Warmup-Stable-Decay 14.9 Post-Training: SFT, DPO, and Narrow RLVR (GRPO) That Works at 100M The Math Reference Sheet
optimization (55)
1.1 Linear Algebra for Deep Learning 1.3 Calculus, Optimization & Convexity 1.4 Numerical Computing, Floating Point & Precision 1.5 Machine Learning Fundamentals 1.6 Neural Networks From Scratch: MLPs & Backprop 1.8 GPU Architecture & The Memory Hierarchy 1.9 Parallel Computing & Collective Communication 2.1 Tokenization: BPE, WordPiece, Unigram & Byte-Level 2.11 Beyond Attention: SSMs, Mamba, RWKV & Linear Attention 3.3 The Pretraining Objective & Loss 3.4 Scaling Laws: Kaplan, Chinchilla & Beyond 3.5 Distributed Training I: Data Parallelism, DDP, ZeRO & FSDP 3.7 Megatron-LM, DeepSpeed & Parallelism in Practice 3.8 Mixed Precision, bf16 & FP8 Training 3.9 Optimizers: SGD, Adam, Adafactor, Lion, Muon & Shampoo 3.10 Learning Rate Schedules, Warmup, Batch Size & Hyperparameters 3.11 Training Stability, Loss Spikes & Debugging Large Runs 3.14 Data Mixing, Domain Weighting & Curriculum 3.16 Continual & Domain-Adaptive Pretraining 3.17 Training an LLM From Scratch: The End-to-End Recipe 4.1 The Roofline Model & Performance Engineering 4.2 FlashAttention I: IO-Awareness & The Online Softmax 4.3 FlashAttention 2 & 3: Work Partitioning, Warp Specialization & FP8 4.4 Writing GPU Kernels with Triton 4.5 CUDA Programming Essentials for ML Engineers 4.6 PagedAttention & KV-Cache Memory Management 4.7 Quantization I: Post-Training Quantization (GPTQ, AWQ, SmoothQuant) 4.8 Quantization II: INT4/INT8/FP8, GGUF, bitsandbytes & QAT 4.9 Kernel Fusion, torch.compile, CUDA Graphs & Compilers 4.10 Memory-Efficient Training: Checkpointing, Offloading & LoRA Math 5.3 PEFT I: LoRA, QLoRA, DoRA & The Adapter Family 5.4 PEFT II: Prompt/Prefix Tuning, IA3, Model Merging & Soups 5.6 Policy Gradients & PPO for Language Models 5.7 Direct Preference Optimization & Its Variants 5.8 GRPO, RLOO & Critic-Free RL 5.12 Distillation, Model Compression & Knowledge Transfer 6.9 Advantage Estimation, KL Control & Stability Tricks 6.11 Scaling RL: Throughput, Load Balancing & The Latest Tricks 6.12 RL Data, Curriculum & Replay Management 7.1 The Anatomy of LLM Inference: Prefill, Decode & The KV Cache 7.2 Continuous Batching & Request Scheduling 7.3 vLLM: Architecture, PagedAttention & Internals 7.4 SGLang: RadixAttention & Structured Programs 7.6 Speculative Decoding: Draft Models, Medusa, EAGLE & Lookahead 7.7 Prefix Caching & KV-Cache Reuse 7.12 Inference Economics: Latency, Throughput & Cost 7.14 Multi-Tenant LoRA & Adapter Serving at Scale 8.4 Context Engineering & Management 12.3 Caching, Routing & Cost Control in Production 14.1 The Capstone: Building Stack-100M, and the 2026 Small-Model Landscape 14.5 Mini Scaling Laws: Fit Your Own Law Before Spending the Budget 14.6 Optimizer & Schedule: Muon + MuonClip and Warmup-Stable-Decay 14.7 The Pretraining Run: A Complete Single-GPU Training Loop 14.8 Mid-Training: Quality Annealing, Long-Context Extension & Capability Injection The Math Reference Sheet
production (29)
3.1 Pretraining Data: Sources, Crawling & The Data Pipeline 3.2 Data Cleaning, Deduplication & Quality Filtering 3.12 Checkpointing, Fault Tolerance & Long-Running Jobs 5.4 PEFT II: Prompt/Prefix Tuning, IA3, Model Merging & Soups 7.7 Prefix Caching & KV-Cache Reuse 7.9 Sampling Strategies & Decoding Algorithms 8.3 Harness Engineering: Building a Coding Agent 8.5 Memory Systems for Agents 8.6 The Model Context Protocol (MCP) 8.7 Multi-Agent Systems & Orchestration 8.8 Agent Evaluation & Benchmarks 8.9 Prompt Engineering as Engineering 9.3 Retrieval-Augmented Generation Architectures 9.4 Chunking, Reranking & Hybrid Search 11.1 The Evaluation Problem & Benchmark Landscape 11.2 LLM-as-a-Judge & Automated Evaluation 11.3 Building Eval Harnesses 11.6 Statistical Rigor in Evaluation: Confidence Intervals & Significance 12.1 Designing an LLM Serving System 12.2 Observability, Logging & LLMOps 12.3 Caching, Routing & Cost Control in Production 12.4 Safety, Guardrails & Content Moderation 12.5 Data Flywheels & Continuous Improvement 12.6 Security: Prompt Injection, Jailbreaks & Defenses 12.7 Online Evaluation: A/B Testing, Canaries & Guardrail Metrics 12.8 Reliability Engineering for LLM Systems: SLOs & Incident Response 13.4 Watermarking, Provenance & AI-Content Detection 13.6 AI Governance, Compliance & Regulation 14.12 Retrospective: Cost Accounting, Reproducibility, and the Path to 1B
quantization (11)
1.4 Numerical Computing, Floating Point & Precision 3.8 Mixed Precision, bf16 & FP8 Training 4.3 FlashAttention 2 & 3: Work Partitioning, Warp Specialization & FP8 4.7 Quantization I: Post-Training Quantization (GPTQ, AWQ, SmoothQuant) 4.8 Quantization II: INT4/INT8/FP8, GGUF, bitsandbytes & QAT 4.10 Memory-Efficient Training: Checkpointing, Offloading & LoRA Math 5.3 PEFT I: LoRA, QLoRA, DoRA & The Adapter Family 5.12 Distillation, Model Compression & Knowledge Transfer 7.5 TensorRT-LLM, TGI & Other Serving Stacks 7.12 Inference Economics: Latency, Throughput & Cost 14.11 Evaluation & Serving: Honest Benchmarks, int4 Quantization, and Running on a Laptop
rl (28)
5.5 The RLHF Pipeline & Reward Modeling 5.6 Policy Gradients & PPO for Language Models 5.7 Direct Preference Optimization & Its Variants 5.8 GRPO, RLOO & Critic-Free RL 5.9 RL with Verifiable Rewards (RLVR) & The Reasoning Recipe 5.10 Reasoning, Chain-of-Thought & Test-Time Compute 5.11 Constitutional AI, RLAIF & Self-Improvement 5.13 Reward Hacking, Over-Optimization & Alignment Failures 6.1 The Anatomy of an RL-for-LLM System 6.2 The Generation–Training Loop & Rollout Engines 6.3 TRL: HuggingFace's RL Library 6.4 veRL: HybridFlow & The Single-Controller Architecture 6.5 OpenRLHF, NeMo-Aligner & Ray-Based Systems 6.6 Prime-RL, Async RL & Decentralized Training 6.7 Colocated vs Disaggregated RL & Weight Synchronization 6.8 Reward Engineering, Verifiers & Sandboxes 6.9 Advantage Estimation, KL Control & Stability Tricks 6.10 Agentic & Multi-Turn RL 6.11 Scaling RL: Throughput, Load Balancing & The Latest Tricks 6.12 RL Data, Curriculum & Replay Management 8.2 The Agentic Loop: ReAct, Plan-Execute & Reflection 8.7 Multi-Agent Systems & Orchestration 11.2 LLM-as-a-Judge & Automated Evaluation 11.4 Reasoning, Coding & Agentic Evals 14.9 Post-Training: SFT, DPO, and Narrow RLVR (GRPO) That Works at 100M 14.10 A Narrow Auto-Research Agent: ReAct, Tool-Use & Retrieval by Distillation The Math Reference Sheet Key Papers: An Annotated Reading List
safety (18)
5.1 Supervised Fine-Tuning & Instruction Tuning 5.5 The RLHF Pipeline & Reward Modeling 5.9 RL with Verifiable Rewards (RLVR) & The Reasoning Recipe 5.11 Constitutional AI, RLAIF & Self-Improvement 5.13 Reward Hacking, Over-Optimization & Alignment Failures 6.8 Reward Engineering, Verifiers & Sandboxes 8.3 Harness Engineering: Building a Coding Agent 8.6 The Model Context Protocol (MCP) 11.5 Red-Teaming, Safety & Robustness Evaluation 12.4 Safety, Guardrails & Content Moderation 12.6 Security: Prompt Injection, Jailbreaks & Defenses 12.7 Online Evaluation: A/B Testing, Canaries & Guardrail Metrics 13.1 Mechanistic Interpretability & Model Internals 13.2 Knowledge Editing & Machine Unlearning 13.3 Privacy, Memorization & Differential Privacy for LLMs 13.4 Watermarking, Provenance & AI-Content Detection 13.5 AI Safety: Scalable Oversight, Dangerous-Capability Evals & Frontier Safety 13.6 AI Governance, Compliance & Regulation
scaling (35)
1.9 Parallel Computing & Collective Communication 1.10 The Accelerator Landscape: TPUs, Trainium, AMD/ROCm & Gaudi 2.5 Positional Encodings: Sinusoidal, Learned, RoPE & ALiBi 2.9 Mixture-of-Experts (MoE) Architectures 3.1 Pretraining Data: Sources, Crawling & The Data Pipeline 3.4 Scaling Laws: Kaplan, Chinchilla & Beyond 3.5 Distributed Training I: Data Parallelism, DDP, ZeRO & FSDP 3.6 Distributed Training II: Tensor, Pipeline, Sequence & Expert Parallelism 3.7 Megatron-LM, DeepSpeed & Parallelism in Practice 3.10 Learning Rate Schedules, Warmup, Batch Size & Hyperparameters 3.11 Training Stability, Loss Spikes & Debugging Large Runs 3.12 Checkpointing, Fault Tolerance & Long-Running Jobs 3.13 Long-Context Pretraining & Context Extension 3.14 Data Mixing, Domain Weighting & Curriculum 3.16 Continual & Domain-Adaptive Pretraining 3.17 Training an LLM From Scratch: The End-to-End Recipe 5.10 Reasoning, Chain-of-Thought & Test-Time Compute 6.1 The Anatomy of an RL-for-LLM System 6.4 veRL: HybridFlow & The Single-Controller Architecture 6.5 OpenRLHF, NeMo-Aligner & Ray-Based Systems 6.6 Prime-RL, Async RL & Decentralized Training 6.11 Scaling RL: Throughput, Load Balancing & The Latest Tricks 7.11 Multi-GPU & Multi-Node Inference 7.13 Serving Mixture-of-Experts: Expert Parallelism & All-to-All Inference 7.14 Multi-Tenant LoRA & Adapter Serving at Scale 10.1 Vision Transformers & Image Encoders 12.1 Designing an LLM Serving System 14.1 The Capstone: Building Stack-100M, and the 2026 Small-Model Landscape 14.2 Data: Sourcing, Filtering, Dedup, Tokenize & Pack ~20B Tokens 14.3 A Byte-Level BPE Tokenizer From Scratch (and Why Vocab Size Is a Design Lever at 100M) 14.4 The Stack-100M Architecture: SOTA Components, Cited and Assembled 14.5 Mini Scaling Laws: Fit Your Own Law Before Spending the Budget 14.8 Mid-Training: Quality Annealing, Long-Context Extension & Capability Injection 14.12 Retrospective: Cost Accounting, Reproducibility, and the Path to 1B The Math Reference Sheet
serving (28)
2.4 Multi-Head Attention, MQA, GQA & MLA 4.6 PagedAttention & KV-Cache Memory Management 5.3 PEFT I: LoRA, QLoRA, DoRA & The Adapter Family 6.7 Colocated vs Disaggregated RL & Weight Synchronization 7.1 The Anatomy of LLM Inference: Prefill, Decode & The KV Cache 7.2 Continuous Batching & Request Scheduling 7.3 vLLM: Architecture, PagedAttention & Internals 7.4 SGLang: RadixAttention & Structured Programs 7.5 TensorRT-LLM, TGI & Other Serving Stacks 7.6 Speculative Decoding: Draft Models, Medusa, EAGLE & Lookahead 7.7 Prefix Caching & KV-Cache Reuse 7.8 Disaggregated Prefill/Decode & Chunked Prefill 7.9 Sampling Strategies & Decoding Algorithms 7.10 Structured & Constrained Generation 7.11 Multi-GPU & Multi-Node Inference 7.12 Inference Economics: Latency, Throughput & Cost 7.13 Serving Mixture-of-Experts: Expert Parallelism & All-to-All Inference 7.14 Multi-Tenant LoRA & Adapter Serving at Scale 8.4 Context Engineering & Management 9.2 Vector Databases & Approximate Nearest Neighbor Search 10.3 Audio, Speech & Multimodal Fusion 12.1 Designing an LLM Serving System 12.2 Observability, Logging & LLMOps 12.3 Caching, Routing & Cost Control in Production 12.4 Safety, Guardrails & Content Moderation 12.8 Reliability Engineering for LLM Systems: SLOs & Incident Response 14.11 Evaluation & Serving: Honest Benchmarks, int4 Quantization, and Running on a Laptop Tooling & Environment Setup Cheatsheet
systems (90)
1.4 Numerical Computing, Floating Point & Precision 1.7 Automatic Differentiation & PyTorch Internals 1.8 GPU Architecture & The Memory Hierarchy 1.9 Parallel Computing & Collective Communication 1.10 The Accelerator Landscape: TPUs, Trainium, AMD/ROCm & Gaudi 2.1 Tokenization: BPE, WordPiece, Unigram & Byte-Level 2.2 Embeddings & The Input Pipeline 2.3 The Attention Mechanism From Scratch 2.4 Multi-Head Attention, MQA, GQA & MLA 2.7 Building a GPT From Scratch (nanoGPT-style) 2.9 Mixture-of-Experts (MoE) Architectures 2.10 Modern Architecture Improvements & Design Choices 2.11 Beyond Attention: SSMs, Mamba, RWKV & Linear Attention 2.12 Diffusion & Non-Autoregressive Language Models 3.1 Pretraining Data: Sources, Crawling & The Data Pipeline 3.2 Data Cleaning, Deduplication & Quality Filtering 3.5 Distributed Training I: Data Parallelism, DDP, ZeRO & FSDP 3.6 Distributed Training II: Tensor, Pipeline, Sequence & Expert Parallelism 3.7 Megatron-LM, DeepSpeed & Parallelism in Practice 3.8 Mixed Precision, bf16 & FP8 Training 3.9 Optimizers: SGD, Adam, Adafactor, Lion, Muon & Shampoo 3.10 Learning Rate Schedules, Warmup, Batch Size & Hyperparameters 3.11 Training Stability, Loss Spikes & Debugging Large Runs 3.12 Checkpointing, Fault Tolerance & Long-Running Jobs 3.13 Long-Context Pretraining & Context Extension 3.17 Training an LLM From Scratch: The End-to-End Recipe 4.1 The Roofline Model & Performance Engineering 4.2 FlashAttention I: IO-Awareness & The Online Softmax 4.3 FlashAttention 2 & 3: Work Partitioning, Warp Specialization & FP8 4.4 Writing GPU Kernels with Triton 4.5 CUDA Programming Essentials for ML Engineers 4.6 PagedAttention & KV-Cache Memory Management 4.7 Quantization I: Post-Training Quantization (GPTQ, AWQ, SmoothQuant) 4.8 Quantization II: INT4/INT8/FP8, GGUF, bitsandbytes & QAT 4.9 Kernel Fusion, torch.compile, CUDA Graphs & Compilers 4.10 Memory-Efficient Training: Checkpointing, Offloading & LoRA Math 5.2 Chat Templates, Data Formatting & Sequence Packing 5.3 PEFT I: LoRA, QLoRA, DoRA & The Adapter Family 5.12 Distillation, Model Compression & Knowledge Transfer 6.1 The Anatomy of an RL-for-LLM System 6.2 The Generation–Training Loop & Rollout Engines 6.3 TRL: HuggingFace's RL Library 6.4 veRL: HybridFlow & The Single-Controller Architecture 6.5 OpenRLHF, NeMo-Aligner & Ray-Based Systems 6.6 Prime-RL, Async RL & Decentralized Training 6.7 Colocated vs Disaggregated RL & Weight Synchronization 6.8 Reward Engineering, Verifiers & Sandboxes 6.10 Agentic & Multi-Turn RL 6.11 Scaling RL: Throughput, Load Balancing & The Latest Tricks 6.12 RL Data, Curriculum & Replay Management 7.1 The Anatomy of LLM Inference: Prefill, Decode & The KV Cache 7.2 Continuous Batching & Request Scheduling 7.3 vLLM: Architecture, PagedAttention & Internals 7.4 SGLang: RadixAttention & Structured Programs 7.5 TensorRT-LLM, TGI & Other Serving Stacks 7.6 Speculative Decoding: Draft Models, Medusa, EAGLE & Lookahead 7.7 Prefix Caching & KV-Cache Reuse 7.8 Disaggregated Prefill/Decode & Chunked Prefill 7.10 Structured & Constrained Generation 7.11 Multi-GPU & Multi-Node Inference 7.12 Inference Economics: Latency, Throughput & Cost 7.13 Serving Mixture-of-Experts: Expert Parallelism & All-to-All Inference 7.14 Multi-Tenant LoRA & Adapter Serving at Scale 8.1 Tool Use & Function Calling 8.3 Harness Engineering: Building a Coding Agent 8.4 Context Engineering & Management 8.5 Memory Systems for Agents 8.6 The Model Context Protocol (MCP) 8.7 Multi-Agent Systems & Orchestration 9.2 Vector Databases & Approximate Nearest Neighbor Search 9.3 Retrieval-Augmented Generation Architectures 9.4 Chunking, Reranking & Hybrid Search 9.5 Advanced RAG: GraphRAG, Agentic RAG & Long-Context vs RAG 9.6 Multimodal & Visual-Document Retrieval: ColPali & Late Interaction 11.3 Building Eval Harnesses 11.4 Reasoning, Coding & Agentic Evals 12.1 Designing an LLM Serving System 12.2 Observability, Logging & LLMOps 12.3 Caching, Routing & Cost Control in Production 12.7 Online Evaluation: A/B Testing, Canaries & Guardrail Metrics 12.8 Reliability Engineering for LLM Systems: SLOs & Incident Response 14.1 The Capstone: Building Stack-100M, and the 2026 Small-Model Landscape 14.2 Data: Sourcing, Filtering, Dedup, Tokenize & Pack ~20B Tokens 14.3 A Byte-Level BPE Tokenizer From Scratch (and Why Vocab Size Is a Design Lever at 100M) 14.6 Optimizer & Schedule: Muon + MuonClip and Warmup-Stable-Decay 14.7 The Pretraining Run: A Complete Single-GPU Training Loop 14.11 Evaluation & Serving: Honest Benchmarks, int4 Quantization, and Running on a Laptop 14.12 Retrospective: Cost Accounting, Reproducibility, and the Path to 1B Glossary of Terms Tooling & Environment Setup Cheatsheet
training (79)
1.1 Linear Algebra for Deep Learning 1.2 Probability, Statistics & Information Theory 1.3 Calculus, Optimization & Convexity 1.4 Numerical Computing, Floating Point & Precision 1.5 Machine Learning Fundamentals 1.6 Neural Networks From Scratch: MLPs & Backprop 1.7 Automatic Differentiation & PyTorch Internals 1.9 Parallel Computing & Collective Communication 2.2 Embeddings & The Input Pipeline 2.6 The Transformer Block: Norms, Residuals, MLPs & Activations 2.7 Building a GPT From Scratch (nanoGPT-style) 2.8 Architecture Variants: Encoder-Decoder, Decoder-Only & Prefix-LM 2.9 Mixture-of-Experts (MoE) Architectures 2.10 Modern Architecture Improvements & Design Choices 2.12 Diffusion & Non-Autoregressive Language Models 3.3 The Pretraining Objective & Loss 3.4 Scaling Laws: Kaplan, Chinchilla & Beyond 3.5 Distributed Training I: Data Parallelism, DDP, ZeRO & FSDP 3.6 Distributed Training II: Tensor, Pipeline, Sequence & Expert Parallelism 3.7 Megatron-LM, DeepSpeed & Parallelism in Practice 3.8 Mixed Precision, bf16 & FP8 Training 3.9 Optimizers: SGD, Adam, Adafactor, Lion, Muon & Shampoo 3.10 Learning Rate Schedules, Warmup, Batch Size & Hyperparameters 3.11 Training Stability, Loss Spikes & Debugging Large Runs 3.12 Checkpointing, Fault Tolerance & Long-Running Jobs 3.13 Long-Context Pretraining & Context Extension 3.14 Data Mixing, Domain Weighting & Curriculum 3.15 Synthetic Data for Pre- and Post-Training 3.16 Continual & Domain-Adaptive Pretraining 3.17 Training an LLM From Scratch: The End-to-End Recipe 4.1 The Roofline Model & Performance Engineering 4.8 Quantization II: INT4/INT8/FP8, GGUF, bitsandbytes & QAT 4.9 Kernel Fusion, torch.compile, CUDA Graphs & Compilers 4.10 Memory-Efficient Training: Checkpointing, Offloading & LoRA Math 5.1 Supervised Fine-Tuning & Instruction Tuning 5.2 Chat Templates, Data Formatting & Sequence Packing 5.3 PEFT I: LoRA, QLoRA, DoRA & The Adapter Family 5.4 PEFT II: Prompt/Prefix Tuning, IA3, Model Merging & Soups 5.5 The RLHF Pipeline & Reward Modeling 5.6 Policy Gradients & PPO for Language Models 5.7 Direct Preference Optimization & Its Variants 5.8 GRPO, RLOO & Critic-Free RL 5.9 RL with Verifiable Rewards (RLVR) & The Reasoning Recipe 5.10 Reasoning, Chain-of-Thought & Test-Time Compute 5.11 Constitutional AI, RLAIF & Self-Improvement 5.12 Distillation, Model Compression & Knowledge Transfer 6.1 The Anatomy of an RL-for-LLM System 6.2 The Generation–Training Loop & Rollout Engines 6.3 TRL: HuggingFace's RL Library 6.4 veRL: HybridFlow & The Single-Controller Architecture 6.6 Prime-RL, Async RL & Decentralized Training 6.8 Reward Engineering, Verifiers & Sandboxes 6.9 Advantage Estimation, KL Control & Stability Tricks 6.10 Agentic & Multi-Turn RL 6.12 RL Data, Curriculum & Replay Management 8.1 Tool Use & Function Calling 9.1 Embeddings & Representation Learning 10.1 Vision Transformers & Image Encoders 10.2 Vision-Language Models 10.4 Diffusion Models & Generative Modeling (Breadth) 10.5 Unified & Any-to-Any Models 12.5 Data Flywheels & Continuous Improvement 13.2 Knowledge Editing & Machine Unlearning 13.3 Privacy, Memorization & Differential Privacy for LLMs 14.1 The Capstone: Building Stack-100M, and the 2026 Small-Model Landscape 14.2 Data: Sourcing, Filtering, Dedup, Tokenize & Pack ~20B Tokens 14.4 The Stack-100M Architecture: SOTA Components, Cited and Assembled 14.5 Mini Scaling Laws: Fit Your Own Law Before Spending the Budget 14.6 Optimizer & Schedule: Muon + MuonClip and Warmup-Stable-Decay 14.7 The Pretraining Run: A Complete Single-GPU Training Loop 14.8 Mid-Training: Quality Annealing, Long-Context Extension & Capability Injection 14.9 Post-Training: SFT, DPO, and Narrow RLVR (GRPO) That Works at 100M 14.10 A Narrow Auto-Research Agent: ReAct, Tool-Use & Retrieval by Distillation 14.12 Retrospective: Cost Accounting, Reproducibility, and the Path to 1B Glossary of Terms The Math Reference Sheet Key Papers: An Annotated Reading List Tooling & Environment Setup Cheatsheet From-Scratch Code Index
Prerequisite map
Each chapter with what it lets you do and the concepts to know first (linked to where they are taught).
Part I — Mathematical & Systems Foundations
Compute and reason with matmuls, SVD, norms, and the backprop transpose identities behind LoRA.
Prereqs: basic vector and matrix notation derivatives and the chain rule Python and PyTorch tensor basics big-O time-complexity notation
Compute entropy, KL divergence, cross-entropy, and perplexity, and justify cross-entropy as the LLM training loss.
Prereqs:
logarithms and exponentials basic calculus (derivatives, integrals) summation and integral notation the softmax function vectors and basic linear algebra
Analyze loss landscapes via gradients and Hessians, and choose/tune optimizers like SGD and Adam.
Prereqs:
single-variable calculus (derivatives) linear algebra basics (vectors, matrices, eigenvalues) matrix multiplication partial derivatives the concept of a neural network loss function
Reason about floating-point formats and apply stable-arithmetic tricks to keep LLM training numerically sound.
Diagnose bias-variance tradeoffs, apply regularization, and choose correct metrics for any ML model.
Build and hand-verify backprop for an MLP, from scalar autograd to a vectorized NumPy classifier.
Trace PyTorch's autograd tape, write custom Function backwards, and debug memory/precision bugs from strides to NaNs.
Prereqs:
chain rule Jacobian matrix MLP forward and backward pass (backpropagation) ReLU activation function gradient descent basics
Classify any GPU kernel as compute- or memory-bound and reason about occupancy and hardware fit.
Prereqs:
matrix multiplication basic linear algebra floating-point number formats (FP16/BF16/FP32) Python programming basic notion of parallel computing
Reason about GPU collective communication costs and choose the right collective for a parallelism strategy.
Choose and port LLM workloads across TPU, Trainium, AMD ROCm, and Gaudi accelerators using roofline reasoning.
Part II — The Transformer Architecture
Implement and choose a subword tokenizer (BPE, WordPiece, Unigram, byte-level) and reason about its cost/arithmetic tradeoffs.
Prereqs:
basic Python programming UTF-8 and character encoding basics regular expressions the idea of an embedding table big-O complexity basic probability (likelihood, EM at an intuitive level)
Implement and size the token embedding, weight-tied unembedding, and input pipeline for a Transformer LM.
Derive, implement, and mask scaled dot-product attention from scratch in NumPy and PyTorch.
Derive KV-cache size for MHA/MQA/GQA/MLA and implement the MHA-to-GQA conversion in PyTorch.
Implement RoPE/ALiBi positional encodings and extend a model's context length via NTK/YaRN scaling.
Assemble and implement a full pre-norm transformer block: norms, residuals, gated FFN, activations.
Implement, initialize, train, and sample from a complete decoder-only GPT in PyTorch from scratch.
Identify and implement encoder-only, encoder-decoder, decoder-only, and prefix-LM attention masks and their tradeoffs.
Build a sparse MoE layer with routing, load balancing, and capacity, and reason about its systems cost.
Explain and implement the modern transformer recipe (RMSNorm, SwiGLU, RoPE, GQA, QK-norm) used in production LLMs.
Analyze and choose sub-quadratic attention alternatives (SSMs, Mamba, RWKV, RetNet, hybrids) for long-context LLMs.
Build and reason about masked-diffusion LLMs: derive the objective, implement a sampler, and design block diffusion.
Part III — Pretraining at Scale
Build a pretraining data pipeline: source, extract, weight, and shard trillions of tokens.
Prereqs: BPE tokenization next-token prediction training objective basic HTTP/HTML concepts distributed/streaming systems basics (object storage, worker pools)
Build a production pipeline to filter, deduplicate, and decontaminate raw web text for pretraining.
Prereqs:
tokenization web crawl data pipelines (CommonCrawl/WARC/WET) hash functions basic probability and set operations n-gram language models next-token prediction / language modeling objective
Compute, mask, and debug the causal-LM cross-entropy loss, including packing and tokenizer-agnostic BPB scoring.
Derive and apply scaling laws to compute-optimally plan a training run's model size and data.
Implement DDP's bucketed gradient sync and choose/apply ZeRO-1/2/3 or FSDP sharding to fit large models.
Partition huge transformer models via tensor, pipeline, sequence/context, and expert parallelism to train beyond single-GPU memory.
Configure Megatron-LM + DeepSpeed parallelism degrees for a multi-hundred-GPU run and measure hardware efficiency via MFU/HFU.
Implement correct bf16/fp16 AMP training loops and scale FP8 GEMMs for production LLM training.
Pick, derive, and implement an optimizer (AdamW, Adafactor, Lion, Shampoo, Muon) for a memory/speed budget.
Choose, implement, and debug LR schedules, warmup, batch-size scaling, and muP hyperparameter transfer for pretraining.
Diagnose, mitigate, and prevent loss spikes and divergence in large-scale LLM pretraining runs.
Design and implement fault-tolerant, resumable checkpointing for large-scale distributed LLM training.
Extend a pretrained model's context window via RoPE scaling, continued pretraining, and ring attention.
Design, weight, and schedule a pretraining data mixture using epoch budgets, DoReMi, and annealing.
Design synthetic-data pipelines (rephrasing, textbook generation, reasoning distillation) that avoid model collapse.
Adapt a converged base model to new domains/languages/data via re-warm+replay CPT, model growth, and compute planning.
Run a full pretrain pipeline end-to-end, raw corpus to sampled checkpoint, at four hardware scales.
Part IV — Kernels, Efficiency & Quantization
Classify any GPU kernel as compute- or memory-bound and pick the matching optimization.
Derive, implement, and IO-analyze exact FlashAttention via the online softmax and tiling.
Analyze FlashAttention-2/3's work partitioning, warp specialization, and FP8 incoherent-processing tricks.
Write, tile, autotune, and debug fused GPU kernels (softmax, matmul, FlashAttention) in Triton.
Write, profile, and optimize custom CUDA kernels (tiled matmul, fused ops) for LLM workloads.
Compute KV-cache memory, diagnose fragmentation, and implement block-table paging with copy-on-write sharing.
Quantize a trained LLM's weights (and activations) to INT4/INT8 post-hoc using GPTQ, AWQ, or SmoothQuant.
Choose and implement INT4/INT8/FP8/GGUF quantization formats, plus QAT and QLoRA, for efficient deployment.
Fuse kernels, capture CUDA graphs, and use torch.compile to speed up PyTorch models.
Compute a transformer's full training memory budget and apply checkpointing, offloading, and LoRA/QLoRA to fit it on smaller GPUs.
Part V — Post-Training & Alignment
Fine-tune a base model into an instruction-follower via response-only loss masking, with data curation and forgetting mitigation.
Build chat templates, loss masks, and packed batches for training conversational models.
Fine-tune giant LLMs cheaply with LoRA/QLoRA adapters and serve thousands of them per GPU.
Adapt LLMs via soft prompts, prefix tuning, IA3, or by merging fine-tuned checkpoints without retraining.
Train a Bradley-Terry reward model and assemble RLHF's four-model PPO pipeline.
Derive REINFORCE, GAE, and PPO's clipped objective, then implement the RLHF training loop.
Derive DPO's closed-form loss from RLHF, implement it, and pick among IPO/KTO/ORPO/SimPO/CPO variants.
Derive and implement critic-free RL (RLOO, GRPO) and apply the 2025 Dr. GRPO/DAPO fixes.
Build verifiable-reward checkers/sandboxes and explain how correctness-only RL grows emergent reasoning.
Implement and evaluate test-time reasoning strategies from chain-of-thought to MCTS and budget-forced long-thinking models.
Build AI-labeled alignment pipelines (CAI, RLAIF, STaR, ReST, SPIN) that scale beyond human annotation.
Implement knowledge distillation, pruning, and quantization-aware compression to shrink large models.
Diagnose reward hacking via the KL-reward frontier and defend RLHF pipelines with KL control, reward ensembles, and verifiable rewards.
Part VI — RL Infrastructure (Deep Dive)
Map the six components and dataflow of any RL-for-LLM system and diagnose its bottlenecks.
Trace, profile, and optimize the five-phase RL rollout-to-gradient loop and its staleness tradeoffs.
Configure and debug TRL's SFT/Reward/DPO/PPO/GRPO trainers, including PEFT and vLLM integration, for real alignment runs.
Explain and implement veRL's HybridFlow dispatch and 3D-HybridEngine resharding to scale RL-for-LLM training.
Design and choose a Ray/Megatron-based RLHF system, laying out actors and weight sync for a given cluster.
Design async off-policy RL pipelines with staleness corrections and verify untrusted decentralized inference via TOPLOC.
Decide GPU placement (colocated vs disaggregated) and design/reason about RL weight-sync latency.
Design rule-based verifiers, LLM judges, and sandboxes to compute reliable RL training rewards.
Implement and debug production-grade advantage estimators, KL control, and PPO clipping tricks for LLM RL.
Build and train multi-turn agentic RL rollouts with correct action/observation masking and credit assignment.
Diagnose RL generation bubbles and apply oversubscription, partial rollout, and DAPO/Dr. GRPO tricks to fix them.
Select, curriculum-order, and replay RL prompts so every rollout carries gradient signal.
Part VII — Inference & Serving
Estimate LLM serving speed from model size and GPU bandwidth using prefill/decode physics.
Design and implement an iteration-level scheduler that keeps LLM-serving batches pinned near capacity.
Explain and tune vLLM's PagedAttention block manager, scheduler, prefix caching, and multi-LoRA internals.
Explain and implement radix-tree KV-cache reuse and write branching SGLang programs with constrained decoding.
Choose and configure the right serving stack (TensorRT-LLM, TGI, llama.cpp, LMDeploy, MLC-LLM) for a workload.
Derive, implement, and analyze speculative decoding (Medusa, EAGLE, lookahead) to speed up lossless LLM decoding.
Reuse cached KV blocks across requests via content hashing/radix tries to cut TTFT and prefill compute.
Diagnose prefill/decode interference and design chunked or disaggregated serving to fix P99 latency.
Choose and implement decoding algorithms (temperature, top-k/p, min-p, beam search, contrastive decoding) to control LLM output quality.
Implement and reason about FSM/PDA-based logit masking to guarantee LLM output matches a grammar or schema.
Size and configure multi-GPU/multi-node LLM serving using tensor, pipeline, expert, and data parallelism.
Derive $/1M-token costs, size GPU clusters, and pick batching/quantization tradeoffs against latency SLOs.
Design and diagnose large-EP MoE inference: overlap all-to-all, balance hot experts, size dense-equivalent throughput.
Design and build a system serving thousands of LoRA adapters over one shared base model.
Part VIII — Agents & Harness Engineering
Build a robust tool-calling loop: define schemas, parse/validate/dispatch calls, handle errors, run in parallel.
Build a working ReAct agent and apply plan-execute, reflection, and tree-search patterns with guardrails against loops and drift.
Build a coding-agent harness: tools, context assembly, loop, permissions, and verification.
Budget, retrieve, compact, and cache an LLM agent's context window to keep it accurate and cheap.
Architect and implement episodic, semantic, and vector agent memory with compaction and session persistence.
Build, connect, and secure MCP servers exposing tools, resources, and prompts to agent hosts.
Design, implement, and cost-benchmark multi-agent LLM systems across topologies and frameworks.
Select agent benchmarks and compute pass@k, confidence intervals, and harness-adjusted comparisons correctly.
Design, optimize, and evaluate production prompts using CoT, few-shot, DSPy/APE, and caching.
Part IX — Retrieval & RAG
Train and evaluate bi-encoder dense embedding models with contrastive InfoNCE and hard negatives for RAG retrieval.
Build and choose ANN indexes (IVF, PQ, HNSW) to retrieve top-k vectors at scale within a latency budget.
Prereqs: dense vector embeddings cosine similarity and Euclidean distance k-means clustering matrix multiplication / GEMM GPU memory bandwidth vs. compute (memory-bound operations)
Design, build, and evaluate a production RAG pipeline from chunking through RAGAS metrics.
Build a production RAG retrieval pipeline: chunk documents, fuse BM25/dense search, and rerank with cross-encoders.
Diagnose naive RAG failure modes and choose among GraphRAG, agentic, self-correcting, or long-context architectures.
Build OCR-free visual-document retrieval by indexing and scoring page images with ColPali-style late interaction.
Part X — Multimodal & Generative Frontiers
Patchify images into tokens and build ViT-based CLIP, SigLIP, and DINOv2 image encoders.
Design and cost a VLM's vision-to-LLM connector, resolution strategy, and training recipe.
Convert audio to tokens/features and build ASR, TTS, and streaming speech LMs fused with text.
Derive, train, and sample diffusion models (DDPM, DDIM, CFG, latent diffusion, flow matching, diffusion LMs).
Design and train a single transformer that understands and generates text, images, audio, and video.
Part XI — Evaluation
Critically read LLM benchmark claims: spot contamination, saturation, and statistically insignificant score gains.
Build, debias, and calibrate an LLM-as-a-judge pipeline, including Elo-based arena ranking.
Build, run, and statistically validate reproducible eval harnesses like lm-evaluation-harness and HELM.
Prereqs:
next-token prediction and log-probabilities few-shot prompting the benchmark landscape (e.g. MMLU, HellaSwag) tokenization and special/BOS tokens basic statistics (standard error, hypothesis testing) chat/instruction-tuned model formatting
Build trustworthy, execution-verified evals for coding, math reasoning, and long-horizon agentic tasks.
Design, run, and critique safety evals: jailbreak ASR, over-refusal, bias/toxicity, and dangerous-capability elicitation.
Attach confidence intervals, paired significance tests, and power analysis to any eval score.
Prereqs:
benchmark evaluation basics (accuracy, pass/fail scoring) probability and statistics foundations (sampling distributions, standard error, hypothesis testing) binomial distribution LLM-as-a-judge evaluation eval harness mechanics
Part XII — Production, Systems & MLOps
Design and size a multi-replica LLM serving fleet: routing, SLOs, queueing, memory, batching, autoscaling.
Instrument, log, evaluate, and alert on production LLM systems, closing the ship-observe-improve loop.
Design a layered cache-route-batch-spot stack that cuts LLM production API spend 50-90%.
Design, tune, and deploy a layered guardrail stack (input/output classifiers, PII redaction, shield models) around a production LLM.
Build a production data flywheel that logs, labels, distills, retrains, and eval-gates model updates.
Design layered defenses (sandboxing, dual-LLM, structured output) against prompt injection and jailbreaks.
Design A/B tests, canaries, and guardrail metrics to evaluate and safely roll out LLM features in production.
Define quality-aware SLOs, diagnose LLM failures, and build rollback, postmortem, and failover systems.
Part XIII — Interpretability, Safety & Governance
Reverse-engineer transformer internals using probing, logit lens, activation patching, circuits, and SAEs.
Locate and surgically edit facts in model weights, and verify unlearning survives adversarial audits.
Measure and defend against LLM training-data leakage using canaries, MIA audits, and DP-SGD.
Implement, evaluate, and choose among watermarking, C2PA provenance, and post-hoc AI-text detection.
Design scalable-oversight, dangerous-capability, and control protocols to gate frontier model deployment.
Determine which AI-regulation obligations apply to a model and build the compliance artefacts to satisfy them.
Prereqs: training compute (FLOPs) and Chinchilla-style scaling approximations pretraining data pipelines and web crawling red-teaming and safety evaluation methodology watermarking and content-provenance techniques basic notions of model parameters and training tokens
Part XIV — Capstone: Build a 100M LLM End-to-End
Plan and budget a from-scratch 101M-parameter LLM build using 2026 small-model best practices.
Source, filter, deduplicate, tokenize, and pack a 20B-token corpus for a 100M-parameter model.
Build a fast from-scratch byte-level BPE trainer and choose vocab size as a 100M-parameter budget tradeoff.
Assemble and parameter-count a from-scratch 100M-parameter transformer from named, cited 2024-2026 SOTA components.
Fit and validate your own scaling law from a cheap model ladder before committing training budget.
Implement a Muon+AdamW hybrid optimizer with MuonClip/QK-clip and a WSD schedule.
Write and run a complete, resumable single-GPU pretraining loop with real throughput monitoring.
Anneal on premium data, extend RoPE context to 8192, and inject math/code capability across one continuous LR decay.
Post-train a 100M base model with SFT, DPO, and narrow GRPO/RLVR into a stopping, preference-tuned, arithmetic-capable assistant.
Distill teacher ReAct trajectories to SFT a 100M model into a narrow tool-using research agent.
Build an honest small-model eval harness, then quantize and serve the model on a laptop CPU.
Itemize the ~$100 training bill, make a run bit-reproducible, and plan the 100M-to-1B scale-up.
Appendix
Look up a precise, cross-linked definition for any term used anywhere in the LLM stack.
Look up and correctly implement the core LLM-stack equations: norms, attention, RoPE, AdamW, scaling laws, LoRA, and RL losses.
Navigate the field's canonical papers by topic and follow a six-week reading curriculum with textbook cross-links.
Set up, pin, debug, and profile the CUDA/PyTorch/HF/DeepSpeed/vLLM stack for real LLM work.
Locate any from-scratch code implementation in the book and trace what it depends on and enables.