← All writing

Training a transformer on eight GPUs, then serving it

A worked example you can click through. The model has 5 transformer layers with the exact layer shape of Llama 2 7B, 1.27 billion parameters in all. Each training step reads 16 sequences of 2,048 tokens. The machine is one node with 8 GPUs, split three ways: 2‑way tensor parallel × 2‑way pipeline parallel × 2‑way data parallel.

Data parallel ZeRO / FSDP sharding Tensor parallel Pipeline parallel Sequence and context parallel Expert parallel

These colors mean the same thing everywhere below. The page is assembled from published papers and framework documentation up to mid‑2026, without live web access. Numbers are rounded estimates meant for building intuition, and the reading list at the end names the sources so you can check the details.

1

The whole lifecycle

Pretraining gets the headlines, but it is one stage of eight. Data work comes before it, alignment and evaluation after it, and a separate engineering discipline takes over when the weights leave the training cluster. Pick a stage to see what happens there and where this page covers it.

    2

    The example model and batch

    Every number on this page comes from one concrete configuration. It is shallow, but every layer is full size, so the per‑layer arithmetic (shapes, bytes, FLOPs, communication) is what you would see in a real 7B‑class model.

    SettingValueSymbol
    Layers5L
    Hidden size4,096h
    Attention heads32, each of size 128a
    MLP width11,008 (SwiGLU: gate, up and down matrices)f
    Vocabulary32,000 tokensV
    Sequence length2,048 tokenss
    Micro-batch2 sequences = 4,096 tokensb
    Micro-batches per replica4 (gradient accumulation)m
    Data-parallel replicas2d
    Global batch2 × 4 × 2 = 16 sequences = 32,768 tokensB
    Parameters1.012B in the layers (202.4M each) + 0.262B in the embedding and LM head = 1.27BN

    From text to training pairs

    A tokenizer (byte-pair encoding here, with 32,000 entries) turns text into integer ids. Pretraining has a single objective: predict the next token. So the target sequence is the input shifted one position left, and a row of 2,048 tokens yields 2,048 predictions in one forward pass.

    One packed row, cut down to seven tokens

    Input Theid 450 catid 6635 satid 3290 onid 373 theid 278 matid 1775 predictspredictspredictspredictspredictspredicts Target cat sat on the mat.

    Ids are illustrative. Real rows are packed: many documents are concatenated with an end-of-document token between them until the row is exactly 2,048 tokens long, and a document mask stops attention from crossing into a neighbouring document. Nothing is padded, so no compute is wasted.

    3

    One training step on one GPU

    Before splitting anything, look at what a single GPU does with one micro-batch. Pick any block to see its tensor shape for 2 × 2,048 tokens, its parameter count and its cost. Then play the four phases of a step. Chapter 4 then opens up the matrices themselves, number by number.

    The stack, top to bottom

    Inside every layer

    Mixed precision: 16 bytes per parameter

    Matrix multiplies run in bf16 (and increasingly FP8 on Hopper and Blackwell GPUs). The optimizer keeps an fp32 master copy of each weight plus Adam's two running averages, also in fp32. That is 2 (bf16 weight) + 2 (bf16 gradient) + 4 + 4 + 4 = 16 bytes per parameter before a single activation is stored: 20.4 GB for our 1.27B model. Megatron-style frameworks accumulate gradients in fp32, which makes it 18 bytes. bf16 has the same exponent range as fp32, so the loss scaling that fp16 training needed is no longer necessary.

    Gradient accumulation

    A replica is responsible for 8 sequences per step but processes only 2 at a time. Gradients from its 4 micro-batches add into the same buffer, and the optimizer runs once at the end. This separates the batch size the optimizer sees (chosen for training stability, often 4–16M tokens in large runs) from what fits in memory. It is also what feeds the pipeline in chapter 8.

    The optimizer step

    First the gradient is clipped: if the norm of all gradients together exceeds 1.0, every gradient is scaled down to that norm. Then AdamW updates each weight w with gradient g at step t:

    m ← β₁·m + (1 − β₁)·g              β₁ = 0.9
    v ← β₂·v + (1 − β₂)·g²             β₂ = 0.95
    m̂ = m / (1 − β₁ᵗ),  v̂ = v / (1 − β₂ᵗ)
    w ← w − lr_t · ( m̂ / (√v̂ + ε) + λ·w )   ε = 1e-8, weight decay λ = 0.1

    The learning rate lr_t follows a schedule: linear warm-up over the first ~2,000 steps, then cosine decay to about a tenth of the peak. Many recent runs use warmup-stable-decay instead, holding the rate flat and decaying only at the end, which makes it easy to branch off and anneal intermediate checkpoints.

    What a step costs

    A reliable rule: training FLOPs ≈ 6 × parameters × tokens, 2 for the forward pass and 4 for the backward (backward computes two gradients per matmul). For our step, 6 × 1.27B × 32,768 ≈ 250 TFLOP, less about 26 TFLOP for the embedding (a lookup, not a matmul), plus ~16 TFLOP of attention: roughly 240 TFLOP. An H100 peaks near 989 dense bf16 TFLOP/s, and good runs sustain about 40% of that (model FLOPs utilization, MFU). Eight GPUs therefore finish a step in about 0.08 s, around 430,000 tokens per second.

    4

    Follow the matrices

    The figures so far describe tensors by their shapes. This chapter shows the numbers inside them. Real tensors hold millions of entries, so it shrinks the model until every number fits on screen, and then actually runs it: a 12-word vocabulary, 8 hidden dimensions, 2 attention heads of size 4, an MLP 16 wide, and a single layer where the full model stacks five identical ones. The batch is 2 sequences of 7 tokens. Every value is computed live in your browser with the same operations the full model uses: RMSNorm, RoPE, causal attention, SwiGLU, cross-entropy, backpropagation and AdamW.

    Step through from raw text to a generated sentence, and hover or tap any cell to see exactly how it was computed. The model starts with random weights. At step 21 you can train it, then revisit earlier steps to see how every matrix changed. The GPU split menu draws where our 8-GPU layout would cut each matrix.

    Negative value Positive value or probability Inputs used by the hovered cell
    5

    Why one GPU runs out

    An honest note first: our toy fits on one 80 GB GPU. Its 20 GB of model state plus about 3 GB of activations per micro-batch is comfortable, and a real team would train it with plain data parallelism. We split it three ways only so that every mechanism is visible at a size you can reason about. Real models hit two walls.

    Weights, gradients and Adam state at 16 bytes per parameter (log scale)

    Activations come on top of this and grow with sequence length and micro-batch size.

    The memory wall. Llama 3 70B needs 1.13 TB just for weights, gradients and optimizer state, more than fourteen GPUs' worth before a single activation. The compute wall. Llama 3.1 405B trained on about 15.6 trillion tokens: 6 × 405B × 15.6T ≈ 3.8 × 10²⁵ FLOP. One H100 at 400 TFLOP/s would need about 3,000 years. Meta used 16,384 H100s for a couple of months.

    Every technique below attacks one or both walls by splitting something different.

    TechniqueWhat it splitsWhat it savesCommunicationTypical degree
    Data parallelThe batchTime (not memory)All-reduce gradients once per step, overlapped with backward8 to thousands
    ZeRO / FSDPOptimizer state, gradients, weightsModel state ÷ dReduce-scatter + all-gatherSame group as DP
    Tensor parallelEvery weight matrixWeights and activations ÷ tFour all-reduces per layer per micro-batch2–8, inside a node
    Sequence parallelActivations in norms and residualsThe rest of activations ÷ tReplaces each all-reduce with reduce-scatter + all-gatherSame as TP
    Pipeline parallelThe layersWeights ÷ pPoint-to-point sends at stage boundaries2–64, across nodes
    Context parallelThe sequence, all the way throughActivations ÷ cK/V passed around a ring, or all-to-all2–32, for long context
    Expert parallelMixture-of-experts expertsExpert weights ÷ eTwo all-to-alls per MoE layer8–64+
    6

    Data parallelism and ZeRO

    Give each GPU a full copy of the model and a different slice of the batch. After backward, every copy holds a different gradient. Average them, and every copy applies the identical update, so the replicas never drift apart. The result is mathematically the same as one GPU processing all 16 sequences, because the average gradient does not care who computed which part.

    The averaging is an all-reduce. The standard implementation (in NCCL, the library nearly every framework uses) is the ring: split the gradient into N chunks and pass them around a ring of GPUs twice, once to add them up (reduce-scatter) and once to hand the sums around (all-gather). Step through it below.

    Ring all-reduce on 4 GPUs

    Each GPU sends 2 × (N − 1)/N of the gradient size in total, whether N is 4 or 4,000. That is why data parallelism scales so well. Two tricks make it nearly free: bucketing groups gradients into buffers of ~25–100 MB so each collective is large enough to be efficient, and overlap starts all-reducing the last layers' bucket while backward is still computing earlier layers.

    ZeRO and FSDP: stop storing the same thing d times

    Plain data parallelism keeps all 16 bytes per parameter on every GPU. ZeRO (from DeepSpeed) and FSDP (PyTorch's version) shard that state across the data-parallel group instead. Notice in the ring that the reduce-scatter half already leaves each GPU owning the sum for one chunk. ZeRO simply stops there and lets each GPU update only the part it owns.

    Model state per GPU

    bf16 weights (2 B) bf16 gradients (2 B) fp32 master + Adam m, v (12 B)
    • Stage 1 shards the optimizer state. Gradients are reduce-scattered, each GPU runs AdamW on its 1/d, then the updated bf16 weights are all-gathered. The traffic equals one all-reduce, so this is free memory. Megatron calls it the distributed optimizer; our 8-GPU example uses it.
    • Stage 2 also shards gradients: each bucket is reduce-scattered straight into its owner and the rest is freed.
    • Stage 3, which is what FSDP does, also shards the weights. Each layer's weights are all-gathered just before its forward, freed, and gathered again for backward. That costs 1.5× the traffic of plain data parallelism, hidden by prefetching the next layer while the current one computes.

    PyTorch's FSDP2 (the basis of TorchTitan) shards per parameter. Hybrid sharding (HSDP) shards inside a node and replicates across nodes, which keeps the frequent all-gathers on fast NVLink while only the once-per-step gradient reduction crosses the network.

    7

    Tensor and sequence parallelism

    When one layer's weights are too big for a GPU, or you want to split activations too, cut every matrix. The Megatron-LM trick is to pair a column split with a row split, so the GPUs only need to exchange data once per block.

    In the MLP, the gate and up matrices are split by columns: GPU 0 gets hidden units 0–5,503 and GPU 1 gets 5,504–11,007. SwiGLU is element-wise, so each GPU applies it to its own half without talking to the other. The down matrix is split by rows to match, so each GPU produces a partial sum of the full output, and one all-reduce adds the two partials. Attention splits the same way along heads: 16 heads per GPU, then a row-split output projection and one all-reduce.

    One block split across a tensor-parallel pair

    The cost. Each layer needs 2 all-reduces in forward and 2 in backward, each on a [2, 2048, 4096] bf16 tensor of 32 MiB, per micro-batch. That is far more traffic than data parallelism, and it sits on the critical path: the next matmul cannot start until the all-reduce is done. So TP stays inside a node on NVLink (about 450 GB/s each way per H100) and almost never crosses the network (about 50 GB/s per GPU on 400 Gb/s InfiniBand). TP = 8, one full node, is the usual ceiling.

    Sequence parallelism. With TP alone, the RMSNorms and residual adds between the split matmuls run redundantly: both GPUs hold the full [2, 2048, 4096] tensor. Sequence parallelism splits those regions along the sequence instead, so each GPU normalizes 1,024 of the 2,048 positions. The all-reduce becomes a reduce-scatter that leaves each GPU with its share of positions, and the next block begins with an all-gather. An all-reduce is exactly a reduce-scatter followed by an all-gather, so the bytes on the wire do not change, but the activation memory of those regions drops by t.

    Embedding and loss. The vocabulary is split too. Each GPU stores 16,000 embedding rows, looks up only the ids it owns, and an all-reduce fills in the rest. At the output, each GPU computes 16,000 of the 32,000 logits. The cross-entropy then needs only three tiny all-reduces (the per-position maximum, the sum of exponentials, and the target token's logit), so the full 131-million-entry logit tensor never has to exist on one GPU.

    8

    Pipeline parallelism

    Split the layers into stages on different GPUs. Our stage 0 holds the embedding and layers 1–3; stage 1 holds layers 4–5 and the LM head. The split is uneven on purpose: the head's 131M-parameter matmul costs about 60% of a layer, and the loss adds more, so the last stage gets fewer layers. Stages exchange only the activation at the boundary, one [2, 2048, 4096] tensor per micro-batch, which is cheap enough to send across nodes.

    The catch is idle time. Stage 1 cannot start until stage 0 finishes the first micro-batch, and stage 0 sits idle at the end waiting for the last gradient. Feeding the pipeline many micro-batches fills it; the idle time left over is the bubble, a fraction (p − 1)/(m + p − 1) of the step for p stages and m micro-batches.

    Pipeline schedules (forward takes 1 time unit, backward takes 2)

    Stages
    Forward Backward Bubble (idle)

    GPipe runs all forwards and then all backwards, so each stage holds activations for all m micro-batches at once. 1F1B (one forward, one backward, from PipeDream and Megatron) starts each backward as early as possible. The bubble is the same, but stage i never holds more than p − i micro-batches, so memory stops growing with m. Interleaved 1F1B gives each GPU several non-adjacent chunks of layers (2 here: with 2 stages, GPU 0 might run layers 1–2 and 5–6 of a deeper model), which divides the bubble by the number of chunks at the cost of more sends.

    Newer schedules go further. Zero-bubble pipelining splits the backward pass into its two halves, the input gradient (needed right away by the previous stage) and the weight gradient (needed only by the optimizer), and slides the weight half into the gaps. DeepSeek's DualPipe feeds micro-batches from both ends of the pipeline at once and overlaps computation with communication. A rule of thumb for plain 1F1B is to keep m at 4p or more. In our step m = 4 and p = 2, so the bubble is 20%.

    9

    Context parallelism for long sequences

    Activation memory grows in proportion to sequence length. At 128K tokens a single sequence of a large model no longer fits, even with TP = 8. Context parallelism splits the sequence across GPUs for the whole layer. The MLP and norms do not mind, since every position is processed independently. Attention is the one place where a position needs the others, and it needs them only through their keys and values.

    Ring attention: each GPU keeps its own queries while the key/value blocks travel around a ring. After c − 1 hops every query has met every key. FlashAttention's online softmax merges the partial results exactly, and the next block's transfer overlaps with the current block's compute. The grid below shows which tiles of the attention matrix get computed at each hop, for 4 GPUs.

    Causal attention tiles, rows = query chunks, columns = key/value chunks

    Ring hop

    Causal masking makes the naive split unfair: the GPU holding the first chunk has almost nothing to compute, and the one holding the last chunk has the most. The standard fix (used in Megatron and in Llama 3's long-context training) cuts the sequence into 2c chunks and gives GPU i chunks i and 2c − 1 − i, pairing an early chunk with a late one.

    The other main design, DeepSpeed-Ulysses, uses an all-to-all to switch from "each GPU has some positions and all heads" to "each GPU has all positions and some heads" for the attention step, and another all-to-all to switch back. It is simpler, but its degree is capped by the number of KV heads. Llama 3 reportedly went from CP = 1 at 8K context to CP = 16 when extending to 128K.

    10

    Expert parallelism for mixture-of-experts

    A mixture-of-experts (MoE) layer replaces the single MLP with many MLPs, the experts, plus a small router that sends each token to its top-k. Imagine our toy with 8 experts per layer and top-2 routing: MLP parameters grow 8×, but each token only pays for 2. The model would have about 6.0B parameters while spending the compute of about 1.95B per token.

    Expert parallelism puts different experts on different GPUs. Every MoE layer then needs two all-to-alls: dispatch sends each token to the GPUs holding its chosen experts, and combine sends the results back to be weighted by the router's scores.

    16 tokens on 4 GPUs routed to 8 experts (2 per GPU), top-2, capacity 5 tokens per expert

    Load imbalance is the enemy. An expert that gets twice its share makes every GPU wait for it, or overflows its capacity and drops tokens (they skip the MLP and pass through the residual path). Routers are trained with an auxiliary load-balancing loss, or, as in DeepSeek-V3, with a per-expert bias that nudges routing toward under-used experts without an extra loss term. Modern MoEs use many small experts: DeepSeek-V3 has 256 routed experts plus 1 shared expert per layer, with 8 routed experts active per token.

    In practice, attention layers stay data-parallel while the experts are sharded across the same GPUs, so EP usually reuses the data-parallel dimension. The all-to-alls are latency-sensitive, so teams keep EP inside a node where possible and overlap it with computation of another micro-batch.

    11

    Putting it together: one full step on 8 GPUs

    Real runs combine these into a device mesh. Ours is 2 × 2 × 2: each GPU belongs to exactly one tensor-parallel pair, one pipeline and one data-parallel group. Press play to watch a complete step: two replicas each push 4 micro-batches through a 1F1B pipeline, then synchronize gradients, clip, update and share the new weights. Click or drag on the timeline to jump anywhere.

    Forward Backward Data-parallel communication Optimizer TP pair busy talking inside layers

    Where each group sits in a real cluster

    The more often a group talks, the closer its members must sit. Tensor parallelism communicates inside every layer, so it goes innermost, on NVLink within a node. Context and expert parallelism come next. Pipeline parallelism only sends at stage boundaries, so its stages can live on different nodes. Data parallelism talks once per step and overlaps with backward, so it goes outermost and spans the whole cluster. Frameworks build this as an n-dimensional mesh (Megatron's rank ordering, PyTorch's DeviceMesh, JAX's Mesh) and derive every process group from it.

    RunGPUsTPCPPPDP / EPNotes
    Our toy8212DP 2ZeRO-1, 1F1B, 4 micro-batches
    Llama 3.1 405B, 8K context16,384 H1008116DP 128bf16, sharded optimizer and gradients, interleaved pipeline
    Llama 3.1 405B, 128K context16,384 H10081616DP 8Long-context extension phase
    DeepSeek-V3, 671B MoE (37B active)2,048 H8001116EP 64 over 8 nodes, ZeRO-1 DPFP8 training, DualPipe, no tensor parallelism

    Configurations as I recall them from the Llama 3 and DeepSeek-V3 technical reports; check the reports for exact values.

    Memory and speed calculator

    Change the layout and see what lands on the busiest GPU. It uses the formulas from this page: 16 bytes of model state per parameter, sharded as ZeRO dictates; about 34·s·b·h bytes of activations per layer per micro-batch divided by the TP degree (FlashAttention and sequence parallelism on); 1F1B keeping p micro-batches in flight on the first stage; and fp32 logits on the last. Try the 405B preset, then switch activation recompute off.

    Weights Gradients Optimizer state Activations Logits

    Collective operations, one table

    CollectiveResultBytes sent per GPU (ring, tensor size S)Used by
    All-reduceEveryone has the sum2 (N−1)/N × SDP gradients, TP block outputs, gradient norm
    Reduce-scatterEveryone has 1/N of the sum(N−1)/N × SZeRO gradients, sequence parallelism
    All-gatherEveryone has all N pieces(N−1)/N × SFSDP weights, ZeRO updated weights, sequence parallelism
    All-to-allEach GPU sends a different piece to each other GPU(N−1)/N × SMoE dispatch and combine, Ulysses attention
    Send / receiveOne GPU to one GPUSPipeline activations and gradients, ring attention
    BroadcastOne GPU's tensor to everyone≈ SInitial weights, pushing weights to RL rollout engines
    12

    Keeping a long run alive

    A frontier pretraining run lasts weeks to months on thousands of GPUs, and the engineering that keeps it going matters as much as the parallelism.

    • Checkpoints. Every few hundred steps, each rank writes its own shard of weights, optimizer state, data-loader position and random-number state, all in parallel, to distributed storage. Asynchronous checkpointing copies to CPU memory first so training resumes within seconds. Resharding tools can load a checkpoint saved at TP = 8, PP = 16 into a different layout.
    • Failures are routine. Meta reported several hundred job interruptions over a 54-day stretch of Llama 3 405B pretraining, most of them hardware: GPUs, HBM memory, network links. At 16,000 GPUs something breaks every few hours, so restart time counts as much as step time. Teams run health checks between jobs, keep spare nodes ready, and set collective timeouts so a hung GPU is detected in minutes.
    • Loss spikes. Watch the loss, the gradient norm and per-layer activation statistics. Common responses: roll back to an earlier checkpoint and skip the batches that caused the spike, lower the learning rate, and use stabilizers such as QK-norm and a z-loss on the logits.
    • Throughput. MFU (model FLOPs utilization) is the headline number. Dense runs typically land at 35–50%. The rest goes to communication that could not be hidden, pipeline bubbles, memory-bound kernels, and stragglers, the slowest GPU that every collective waits for.
    • Decisions made before the big run. Data mixture, learning rate, batch size and architecture are chosen with ablations on small proxy models and scaling laws fitted to them. The big run should hold no surprises.
    13

    Post-training: from text predictor to assistant

    A pretrained model continues text; it does not follow instructions. Post-training uses a small fraction of pretraining compute and reuses the same training stack and parallelism, but at smaller scale with many more iterations.

    • Supervised fine-tuning (SFT). Same loop as pretraining, but on prompts paired with good responses, written by people or by stronger models. The loss is computed only on response tokens. LoRA, which trains small low-rank adapter matrices, makes cheap variants possible.
    • Preference tuning. Classic RLHF trains a reward model on pairs of responses that people ranked, then optimizes the model against it with PPO, with a KL penalty that keeps it close to the SFT model. DPO skips the reward model and the RL loop: a closed-form loss applied directly to chosen/rejected pairs.
    • RL with verifiable rewards. For math, code and tool use, a checker decides whether an answer is right. GRPO samples a group of answers per prompt and uses each answer's reward relative to the group average as its advantage, so no separate value network is needed. For reasoning models this stage has grown into a large share of total compute.
    • The RL system. RL training is two systems in a loop: a trainer (Megatron or FSDP, with the parallelism above) and a fleet of inference engines (vLLM, SGLang) generating rollouts. After each update the new weights are broadcast to the engines. Generation usually dominates the time, so rollouts are batched heavily and are sometimes allowed to run a step behind the trainer (asynchronous RL).
    14

    Deployment and inference

    From training checkpoint to servable weights

    • Merge the shards. Concatenate tensor-parallel pieces along the dimension they were split on (columns for Q/K/V, gate and up; rows for the output and down projections), stack the pipeline stages back into one list of layers, and drop the optimizer state. Our toy shrinks from 20.4 GB of training state to 2.5 GB of bf16 weights.
    • Convert to the format your engine loads, usually safetensors in the Hugging Face layout, or a compiled TensorRT-LLM engine.
    • Quantize. FP8 weights and activations with per-channel scales are close to lossless on H100-class GPUs. INT4 weight-only quantization (AWQ, GPTQ) quarters memory for bandwidth-bound serving. An FP8 KV cache halves the cache.
    • Validate. Re-run evaluations on the converted, quantized model. Conversion bugs, such as a transposed shard or a wrong RoPE setting, are common and silent.

    Two phases of every request

    Prefill runs the whole prompt through the model in one forward pass, like a training forward pass with no backward. It is compute-bound, and it writes every prompt token's keys and values into the KV cache. Decode then produces one token per forward pass, attending to everything cached. Each decode step reads all the weights and the whole KV cache to do only a couple of FLOPs per byte, so it is bound by memory bandwidth. Hence the two latency numbers everyone tracks: time to first token (TTFT, dominated by prefill) and time per output token (TPOT, set by decode). Steps 22 and 23 of chapter 4 run both phases on real numbers, including the growing KV cache.

    The KV cache holds 2 (K and V) × layers × KV heads × head size × bytes per value for every token. For our toy that is 2 × 5 × 32 × 128 × 2 = 80 KB per token. Grouped-query attention (GQA), where 8 KV heads are shared by all 32 query heads, cuts it to 20 KB; that is why nearly every recent model uses GQA or a compressed variant like DeepSeek's multi-head latent attention.

    KV cache and decode speed estimator (H100 80 GB, 3.35 TB/s)

    Weights KV cache

    Batching is where the money is

    Because decode is bandwidth-bound, serving 64 sequences costs little more per step than serving 1: the weights are read once per step either way. Try it in the estimator: aggregate throughput climbs almost linearly with concurrency until the KV cache fills memory or the step becomes compute-bound. Continuous batching (introduced by Orca, standard since vLLM) adds and removes sequences at every step instead of waiting for a whole batch to finish.

    10 requests, 4 batch slots, each bar is one request's decode steps

    The rest of the serving toolbox

    • PagedAttention. vLLM stores the KV cache in fixed-size blocks, like virtual-memory pages, so sequences of unknown length do not fragment memory. Blocks can be shared, which enables prefix caching: a system prompt shared by many requests is prefilled once.
    • Chunked prefill. Long prompts are split into chunks and mixed into decode batches, so one 100K-token prompt does not freeze every other user's stream.
    • Speculative decoding. A cheap drafter (a small model, extra prediction heads, or the model's own multi-token-prediction layers) proposes several tokens, and the big model verifies them all in one forward pass. The output distribution is unchanged, and decode often runs 2–3× faster.
    • Inference parallelism. Tensor parallelism within a node to cut per-token latency; pipeline parallelism across nodes only for models too big for one node; for MoE, expert parallelism across many GPUs with data-parallel attention; and replicas (plain data parallelism) for throughput.
    • Prefill/decode disaggregation. Prefill and decode want different hardware balances, so large deployments run them on separate GPU pools and ship the KV cache between them.

    A production serving stack

    Operating it. Set service-level objectives on TTFT and TPOT percentiles, autoscale on queue depth and KV-cache use, roll new checkpoints out as a canary to a slice of traffic first, and log enough to feed evaluations and abuse detection. Common engines are vLLM, SGLang and TensorRT-LLM, usually on Kubernetes behind a KV-cache-aware router (projects like NVIDIA Dynamo and llm-d). Whatever users teach you becomes data for the next run, and the cycle in chapter 1 starts again.

    15

    Reading list

    Listed from memory, without web access, so titles and years may be slightly off. Search for them rather than trusting every detail here.

    • Shoeybi et al., Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism (2019): tensor parallelism.
    • Narayanan et al., Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM (2021): combining TP, PP and DP; interleaved 1F1B.
    • Korthikanti et al., Reducing Activation Recomputation in Large Transformer Models (2022): sequence parallelism and the activation-memory formula used here.
    • Rajbhandari et al., ZeRO: Memory Optimizations Toward Training Trillion Parameter Models (2020); Zhao et al., PyTorch FSDP (2023).
    • Huang et al., GPipe (2019); Narayanan et al., PipeDream (2019); Qi et al., Zero Bubble Pipeline Parallelism (2023).
    • Liu et al., Ring Attention with Blockwise Transformers (2023); Jacobs et al., DeepSpeed Ulysses (2023).
    • Lepikhin et al., GShard (2020); Fedus et al., Switch Transformers (2021).
    • Dao et al., FlashAttention (2022) and its sequels.
    • Meta, The Llama 3 Herd of Models (2024); DeepSeek-AI, DeepSeek-V3 Technical Report (2024).
    • Ouyang et al., Training language models to follow instructions with human feedback (2022); Rafailov et al., Direct Preference Optimization (2023); Shao et al., DeepSeekMath (2024), which introduced GRPO.
    • Yu et al., Orca (2022); Kwon et al., Efficient Memory Management for LLM Serving with PagedAttention (2023); Leviathan et al., Fast Inference from Transformers via Speculative Decoding (2023); Zhong et al., DistServe (2024).
    • Hugging Face, The Ultra-Scale Playbook (2025): a long, practical walkthrough of the same ground.