An open, ground-up field guide
The LLM StackFrom Silicon to Agents
A ground-up field guide to building, training, aligning, serving, and reasoning with large language models.
14
Parts
146
Chapters
~3,773
Pages
1132k
Words
Start reading →Learning mapGlossaryTools & calculatorsInterview companionPrint / PDF edition
IPart I — Mathematical & Systems Foundations
- Linear Algebra for Deep Learning
- Probability, Statistics & Information Theory
- Calculus, Optimization & Convexity
- Numerical Computing, Floating Point & Precision
- Machine Learning Fundamentals
- Neural Networks From Scratch: MLPs & Backprop
- Automatic Differentiation & PyTorch Internals
- GPU Architecture & The Memory Hierarchy
- Parallel Computing & Collective Communication
- The Accelerator Landscape: TPUs, Trainium, AMD/ROCm & Gaudi
IIPart II — The Transformer Architecture
- Tokenization: BPE, WordPiece, Unigram & Byte-Level
- Embeddings & The Input Pipeline
- The Attention Mechanism From Scratch
- Multi-Head Attention, MQA, GQA & MLA
- Positional Encodings: Sinusoidal, Learned, RoPE & ALiBi
- The Transformer Block: Norms, Residuals, MLPs & Activations
- Building a GPT From Scratch (nanoGPT-style)
- Architecture Variants: Encoder-Decoder, Decoder-Only & Prefix-LM
- Mixture-of-Experts (MoE) Architectures
- Modern Architecture Improvements & Design Choices
- Beyond Attention: SSMs, Mamba, RWKV & Linear Attention
- Diffusion & Non-Autoregressive Language Models
IIIPart III — Pretraining at Scale
- Pretraining Data: Sources, Crawling & The Data Pipeline
- Data Cleaning, Deduplication & Quality Filtering
- The Pretraining Objective & Loss
- Scaling Laws: Kaplan, Chinchilla & Beyond
- Distributed Training I: Data Parallelism, DDP, ZeRO & FSDP
- Distributed Training II: Tensor, Pipeline, Sequence & Expert Parallelism
- Megatron-LM, DeepSpeed & Parallelism in Practice
- Mixed Precision, bf16 & FP8 Training
- Optimizers: SGD, Adam, Adafactor, Lion, Muon & Shampoo
- Learning Rate Schedules, Warmup, Batch Size & Hyperparameters
- Training Stability, Loss Spikes & Debugging Large Runs
- Checkpointing, Fault Tolerance & Long-Running Jobs
- Long-Context Pretraining & Context Extension
- Data Mixing, Domain Weighting & Curriculum
- Synthetic Data for Pre- and Post-Training
- Continual & Domain-Adaptive Pretraining
- Training an LLM From Scratch: The End-to-End Recipe
IVPart IV — Kernels, Efficiency & Quantization
- The Roofline Model & Performance Engineering
- FlashAttention I: IO-Awareness & The Online Softmax
- FlashAttention 2 & 3: Work Partitioning, Warp Specialization & FP8
- Writing GPU Kernels with Triton
- CUDA Programming Essentials for ML Engineers
- PagedAttention & KV-Cache Memory Management
- Quantization I: Post-Training Quantization (GPTQ, AWQ, SmoothQuant)
- Quantization II: INT4/INT8/FP8, GGUF, bitsandbytes & QAT
- Kernel Fusion, torch.compile, CUDA Graphs & Compilers
- Memory-Efficient Training: Checkpointing, Offloading & LoRA Math
VPart V — Post-Training & Alignment
- Supervised Fine-Tuning & Instruction Tuning
- Chat Templates, Data Formatting & Sequence Packing
- PEFT I: LoRA, QLoRA, DoRA & The Adapter Family
- PEFT II: Prompt/Prefix Tuning, IA3, Model Merging & Soups
- The RLHF Pipeline & Reward Modeling
- Policy Gradients & PPO for Language Models
- Direct Preference Optimization & Its Variants
- GRPO, RLOO & Critic-Free RL
- RL with Verifiable Rewards (RLVR) & The Reasoning Recipe
- Reasoning, Chain-of-Thought & Test-Time Compute
- Constitutional AI, RLAIF & Self-Improvement
- Distillation, Model Compression & Knowledge Transfer
- Reward Hacking, Over-Optimization & Alignment Failures
VIPart VI — RL Infrastructure (Deep Dive)
- The Anatomy of an RL-for-LLM System
- The Generation–Training Loop & Rollout Engines
- TRL: HuggingFace's RL Library
- veRL: HybridFlow & The Single-Controller Architecture
- OpenRLHF, NeMo-Aligner & Ray-Based Systems
- Prime-RL, Async RL & Decentralized Training
- Colocated vs Disaggregated RL & Weight Synchronization
- Reward Engineering, Verifiers & Sandboxes
- Advantage Estimation, KL Control & Stability Tricks
- Agentic & Multi-Turn RL
- Scaling RL: Throughput, Load Balancing & The Latest Tricks
- RL Data, Curriculum & Replay Management
VIIPart VII — Inference & Serving
- The Anatomy of LLM Inference: Prefill, Decode & The KV Cache
- Continuous Batching & Request Scheduling
- vLLM: Architecture, PagedAttention & Internals
- SGLang: RadixAttention & Structured Programs
- TensorRT-LLM, TGI & Other Serving Stacks
- Speculative Decoding: Draft Models, Medusa, EAGLE & Lookahead
- Prefix Caching & KV-Cache Reuse
- Disaggregated Prefill/Decode & Chunked Prefill
- Sampling Strategies & Decoding Algorithms
- Structured & Constrained Generation
- Multi-GPU & Multi-Node Inference
- Inference Economics: Latency, Throughput & Cost
- Serving Mixture-of-Experts: Expert Parallelism & All-to-All Inference
- Multi-Tenant LoRA & Adapter Serving at Scale
VIIIPart VIII — Agents & Harness Engineering
- Tool Use & Function Calling
- The Agentic Loop: ReAct, Plan-Execute & Reflection
- Harness Engineering: Building a Coding Agent
- Context Engineering & Management
- Memory Systems for Agents
- The Model Context Protocol (MCP)
- Multi-Agent Systems & Orchestration
- Agent Evaluation & Benchmarks
- Prompt Engineering as Engineering
IXPart IX — Retrieval & RAG
XPart X — Multimodal & Generative Frontiers
XIPart XI — Evaluation
XIIPart XII — Production, Systems & MLOps
- Designing an LLM Serving System
- Observability, Logging & LLMOps
- Caching, Routing & Cost Control in Production
- Safety, Guardrails & Content Moderation
- Data Flywheels & Continuous Improvement
- Security: Prompt Injection, Jailbreaks & Defenses
- Online Evaluation: A/B Testing, Canaries & Guardrail Metrics
- Reliability Engineering for LLM Systems: SLOs & Incident Response
XIIIPart XIII — Interpretability, Safety & Governance
XIVPart XIV — Capstone: Build a 100M LLM End-to-End
- The Capstone: Building Stack-100M, and the 2026 Small-Model Landscape
- Data: Sourcing, Filtering, Dedup, Tokenize & Pack ~20B Tokens
- A Byte-Level BPE Tokenizer From Scratch (and Why Vocab Size Is a Design Lever at 100M)
- The Stack-100M Architecture: SOTA Components, Cited and Assembled
- Mini Scaling Laws: Fit Your Own Law Before Spending the Budget
- Optimizer & Schedule: Muon + MuonClip and Warmup-Stable-Decay
- The Pretraining Run: A Complete Single-GPU Training Loop
- Mid-Training: Quality Annealing, Long-Context Extension & Capability Injection
- Post-Training: SFT, DPO, and Narrow RLVR (GRPO) That Works at 100M
- A Narrow Auto-Research Agent: ReAct, Tool-Use & Retrieval by Distillation
- Evaluation & Serving: Honest Benchmarks, int4 Quantization, and Running on a Laptop
- Retrospective: Cost Accounting, Reproducibility, and the Path to 1B
Cite this book
@book{kagitha_llm_stack_2026,
title = {The LLM Stack: From Silicon to Agents},
author = {Kagitha, Prakash},
year = {2026},
url = {https://prakashkagitha.github.io/llm-stack-book/},
note = {Open web textbook, CC BY 4.0}
}