feat(phase-17/17): disaggregated prefill/decode - NVIDIA Dynamo and llm-d

This commit is contained in:
Rohit Ghumare
2026-04-24 12:19:24 +01:00
parent 1d9cf7c832
commit e141506110
5 changed files with 301 additions and 0 deletions
@@ -0,0 +1,69 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 960 460" font-family="Georgia, 'Times New Roman', serif">
<defs>
<style>
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
.pre { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
.dec { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
.router { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
.head { font-size: 12px; font-weight: 700; fill: #1a1a1a; }
.step { font-size: 12px; font-family: 'Menlo', monospace; fill: #222; }
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
.caption { font-size: 11px; fill: #555; font-style: italic; }
.arrow { stroke: #1a1a1a; stroke-width: 1.5; fill: none; marker-end: url(#arr); }
</style>
<marker id="arr" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="5" markerHeight="5" orient="auto">
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
</marker>
</defs>
<text x="480" y="24" text-anchor="middle" class="title">disaggregated prefill + decode — NVIDIA Dynamo / llm-d</text>
<rect x="40" y="60" width="140" height="70" class="router"/>
<text x="110" y="82" text-anchor="middle" class="head">router</text>
<text x="110" y="100" text-anchor="middle" class="small">cache-aware</text>
<text x="110" y="116" text-anchor="middle" class="small">+ SLA planner</text>
<rect x="240" y="60" width="260" height="120" class="pre"/>
<text x="370" y="82" text-anchor="middle" class="head">prefill pool — compute-bound</text>
<text x="370" y="102" text-anchor="middle" class="step">H100 / B200</text>
<text x="250" y="122" class="small">· matmul-heavy forward</text>
<text x="250" y="138" class="small">· FLOPs-limited</text>
<text x="250" y="154" class="small">· scale on queue depth</text>
<text x="370" y="174" text-anchor="middle" class="caption">~2000 TFLOPS FP8 useful</text>
<rect x="640" y="60" width="260" height="120" class="dec"/>
<text x="770" y="82" text-anchor="middle" class="head">decode pool — memory-bound</text>
<text x="770" y="102" text-anchor="middle" class="step">H200 or aggressive quant</text>
<text x="650" y="122" class="small">· one token per iter, all weights</text>
<text x="650" y="138" class="small">· HBM-bandwidth-limited</text>
<text x="650" y="154" class="small">· scale on KV utilization</text>
<text x="770" y="174" text-anchor="middle" class="caption">~3 TB/s HBM3 ceiling</text>
<path class="arrow" d="M180 95 L 235 95"/>
<text x="210" y="88" text-anchor="middle" class="small">prompt</text>
<path class="arrow" d="M500 95 L 635 95"/>
<text x="568" y="78" text-anchor="middle" class="step">NIXL</text>
<text x="568" y="108" text-anchor="middle" class="small">KV transfer</text>
<text x="568" y="125" text-anchor="middle" class="small">RDMA or TCP</text>
<rect x="40" y="210" width="440" height="110" class="box"/>
<text x="260" y="232" text-anchor="middle" class="head">NVIDIA Dynamo</text>
<text x="60" y="256" class="small">· sits above vLLM / SGLang / TRT-LLM</text>
<text x="60" y="274" class="small">· Planner Profiler + SLA Planner auto-configs</text>
<text x="60" y="292" class="small">· Rust core, Python extensibility</text>
<text x="60" y="310" class="small">· 30x on DeepSeek-R1; 50x MoE on GB300 NVL72</text>
<rect x="500" y="210" width="420" height="110" class="box"/>
<text x="710" y="232" text-anchor="middle" class="head">llm-d (Red Hat + AWS)</text>
<text x="520" y="256" class="small">· Kubernetes-native Services per role</text>
<text x="520" y="274" class="small">· packDomain: rack for KV locality</text>
<text x="520" y="292" class="small">· per-role HPA (queue / KV util)</text>
<text x="520" y="310" class="small">· 0.5: hierarchical KV, LoRA routing, UCCL</text>
<rect x="40" y="340" width="880" height="110" class="box"/>
<text x="480" y="362" text-anchor="middle" class="head">when it pays off</text>
<text x="480" y="386" text-anchor="middle" class="step">prompts > 512 tokens AND outputs > 200 tokens</text>
<text x="480" y="404" text-anchor="middle" class="step">MoE serving (DeepSeek-V3, future GPT-5 variants) — double win on expert routing</text>
<text x="480" y="424" text-anchor="middle" class="step">real case: $2M → $1.2M/yr on same workload, same SLA, no new hardware</text>
<text x="480" y="444" text-anchor="middle" class="caption">short prompts: transfer tax dominates, do not disaggregate</text>
</svg>

After

Width:  |  Height:  |  Size: 4.5 KiB

@@ -0,0 +1,59 @@
"""Colocated vs disaggregated serving simulator — stdlib Python.
Models one request through colocated (same GPU) vs disaggregated (prefill pool + decode pool + KV transfer).
Sweeps prompt length to find the crossover.
"""
from __future__ import annotations
# illustrative 2026 constants for 70B FP8 on H100 class
PREFILL_TOK_PER_MS = 40.0 # prefill throughput per GPU per ms
DECODE_TOK_PER_MS_COLOCATED = 0.10
DECODE_TOK_PER_MS_DECODE_GPU = 0.18 # memory-optimized pool (H200-like)
KV_BYTES_PER_TOKEN_70B_FP8 = 125_000
NIXL_RDMA_GB_S = 100
NIXL_TCP_GB_S = 10
def ms_colocated(prompt: int, output: int) -> float:
prefill_ms = prompt / PREFILL_TOK_PER_MS
decode_ms = output / DECODE_TOK_PER_MS_COLOCATED
return prefill_ms + decode_ms
def ms_disaggregated(prompt: int, output: int, use_rdma: bool = True) -> float:
prefill_ms = prompt / PREFILL_TOK_PER_MS
kv_bytes = prompt * KV_BYTES_PER_TOKEN_70B_FP8
transport = NIXL_RDMA_GB_S if use_rdma else NIXL_TCP_GB_S
transfer_ms = (kv_bytes / 1e9) / transport * 1000
decode_ms = output / DECODE_TOK_PER_MS_DECODE_GPU
return prefill_ms + transfer_ms + decode_ms
def main() -> None:
print("=" * 95)
print("DISAGGREGATED vs COLOCATED — same request, different GPU placement")
print("=" * 95)
header = f"{'prompt':>7} {'output':>7} {'colocated (ms)':>15} {'disagg RDMA (ms)':>17} {'disagg TCP (ms)':>16} Winner"
print(header)
print("-" * len(header))
cases = [
(256, 100), (512, 200), (1024, 300), (2048, 400),
(4096, 500), (8192, 800), (16384, 1200), (32768, 2000),
]
for prompt, output in cases:
colo = ms_colocated(prompt, output)
rdma = ms_disaggregated(prompt, output, use_rdma=True)
tcp = ms_disaggregated(prompt, output, use_rdma=False)
winner = "colocated" if colo < rdma else "disaggregated"
print(f"{prompt:>7} {output:>7} {colo:>14.1f} {rdma:>17.1f} {tcp:>16.1f} {winner}")
print()
print("Read: disaggregation wins at longer prompts where decode throughput improvement")
print("on memory-optimized pool outweighs the KV transfer tax. TCP transport raises the")
print("break-even; RDMA makes disaggregation profitable earlier.")
if __name__ == "__main__":
main()
@@ -0,0 +1,142 @@
# Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d
> Prefill is compute-bound; decode is memory-bound. Running both on the same GPU wastes one resource. Disaggregation splits them onto separate pools and transfers KV cache between them over NIXL (RDMA/InfiniBand or TCP fallback). NVIDIA Dynamo (GTC 2025 announce, 1.0 GA) sits above vLLM/SGLang/TRT-LLM — its Planner Profiler + SLA Planner auto-rate-match prefill:decode ratios to meet SLOs. Up to 30x more requests on DeepSeek-R1 on Blackwell with full stack; 50x MoE throughput on GB300 NVL72 + Dynamo. llm-d (Red Hat + AWS) is Kubernetes-native: prefill / decode / router as independent Services with per-role HPA. llm-d 0.5 adds hierarchical KV offloading, cache-aware LoRA routing, UCCL networking, scale-to-zero. Economics: one customer cut $600-800K from a $2M annual inference spend at same request volume, same latency. Short prompts (<512 tokens, short output) don't justify the transfer cost.
**Type:** Learn
**Languages:** Python (stdlib, toy disaggregated-vs-colocated simulator)
**Prerequisites:** Phase 17 · 04 (vLLM Serving Internals), Phase 17 · 08 (Inference Metrics)
**Time:** ~75 minutes
## Learning Objectives
- Explain why prefill and decode have different optimal GPU allocations and quantify the waste under colocation.
- Diagram the disaggregated architecture: prefill pool, decode pool, KV transfer via NIXL, router.
- Name the condition when disaggregation does NOT pay off (short prompts, short outputs).
- Distinguish NVIDIA Dynamo (stack-above) from llm-d (Kubernetes-native) and match each to an operational context.
## The Problem
You run Llama 3.3 70B on 8 H100s. Under mixed workload (long prompts + short outputs), GPUs idle during decode because most of the compute was spent on prefill. Under different workload (short prompts + long outputs), the opposite happens. Colocated prefill + decode means you over-provision both.
Budget impact: 20-40% of GPU time is wasted on the wrong resource. You are buying H100 compute to run memory-bound decode, or buying H100 HBM bandwidth to run compute-bound prefill. Both are expensive waste.
Disaggregation splits prefill and decode onto separate pools sized for each's bottleneck. KV cache transfers from prefill pool to decode pool via high-bandwidth interconnect.
## The Concept
### Why the bottlenecks differ
**Prefill** — run the transformer over the full input prompt in one forward. Matrix multiplications dominate; compute-bound. H100 FP8 gives ~2000 TFLOPS of useful throughput. Batch efficiency is good — one forward processes many tokens.
**Decode** — generate one token at a time, reading the full weights each iteration. Memory-bandwidth-bound. HBM3 gives ~3 TB/s. Batch efficiency is good only at high concurrency — the weights read amortizes across the batch.
Colocating them: you buy GPUs optimized for both. H100 is good at both but costs the same either way. At scale, you want prefill pool on H100 / compute-heavy; decode pool on H200 / memory-heavy, or with aggressive quantization.
### The architecture
```
┌──────────────┐
Request → │ Router │ ───────────────────────┐
└──────┬───────┘ │
│ │
▼ (prompt only) │
┌──────────────┐ KV cache ┌───────▼──────┐
│ Prefill pool │ ─── NIXL ────► │ Decode pool │
│ (compute) │ │ (memory) │
└──────────────┘ └──────┬───────┘
│ tokens
▼
Client
```
NIXL is NVIDIA's inter-node transport. Uses RDMA/InfiniBand when available, TCP fallback otherwise. Transfer latency is real — typically 20-80 ms for KV cache of a 4K-token prompt on 70B FP8. This is why short prompts don't justify disaggregation: the transfer tax exceeds the savings.
### Dynamo vs llm-d
**NVIDIA Dynamo** (GTC 2025 announce, 1.0 GA):
- Sits above vLLM, SGLang, TRT-LLM as an orchestrator.
- Planner Profiler measures workload, SLA Planner auto-configures prefill:decode ratios.
- Rust core, Python extensibility.
- Up to 30x request throughput on DeepSeek-R1 on Blackwell (full stack).
- GB300 NVL72 + Dynamo: 50x MoE throughput vs Hopper.
**llm-d** (Red Hat + AWS, Kubernetes-native):
- Prefill / decode / router as independent Kubernetes Services.
- Per-role HPA with queue depth (prefill) / KV utilization (decode) signals.
- `topologyConstraint packDomain: rack` packs prefill+decode cliques on the same rack for high-bandwidth KV transfer.
- llm-d 0.5 (2026): hierarchical KV offloading, cache-aware LoRA routing, UCCL networking, scale-to-zero.
Use Dynamo if you want a managed stack-above orchestrator. Use llm-d if you want Kubernetes-native primitives and are committed to the CNCF ecosystem.
### Economics
One published case study:
- $2M/year inference spend on colocated serving.
- Switched to disaggregated with Dynamo.
- Same request volume, same P99 latency SLA.
- Savings: $600K-$800K/year (30-40% reduction).
- No new hardware.
The savings come from right-sizing each pool. Prefill-heavy workloads (RAG with 8K+ prefixes) benefit more than balanced.
### When NOT to disaggregate
- Prompts < 512 tokens and outputs < 200 tokens: transfer tax dominates gain.
- Small cluster (< 4 GPUs): not enough pool diversity.
- Team cannot operate two GPU pools with per-role scaling: Dynamo helps but not trivially.
- No RDMA fabric: TCP transfer tax is heavier.
### The router integrates with Phase 17 · 11
Disaggregated routers are KV-cache-aware (Phase 17 · 11). A request lands on the decode pool holding its prefix — if no match, it flows prefill → decode. Hit rate and disaggregation compound — the cache-aware router determines whether a new prefill is even needed.
### MoE on Blackwell is where the real numbers are
GB300 NVL72 + Dynamo shows 50x MoE throughput over Hopper baselines. MoE expert routing is compute-heavy on prefill but memory-heavy on decode (expert caches), so disaggregation is a double win. 2026 frontier model serving is MoE-dominant (DeepSeek-V3, future GPT-5 variants).
### Numbers you should remember
- DeepSeek-R1 on Blackwell + full Dynamo stack: up to 30x request throughput.
- GB300 NVL72 + Dynamo: 50x MoE throughput vs Hopper.
- Real customer case: $600-800K/year savings on $2M spend.
- Disaggregation threshold: prompts >512 tokens + outputs >200 tokens.
- KV transfer via NIXL: 20-80 ms for 4K-prompt KV on 70B FP8.
## Use It
`code/main.py` simulates colocated vs disaggregated serving. Reports throughput, cost per request, and the prompt-length crossover.
## Ship It
This lesson produces `outputs/skill-disaggregation-decider.md`. Given workload and cluster, decides whether to disaggregate.
## Exercises
1. Run `code/main.py`. At what prompt length does disaggregation beat colocation?
2. Design the prefill pool and decode pool for a RAG service with P99 prefix length 8K, output 300.
3. Dynamo vs llm-d: pick one for a pure-Kubernetes shop with no Python runtime preference.
4. Compute KV transfer cost: 4K prefill on 70B FP8 = ~500 MB KV. At RDMA 100 GB/s, transfer = 5 ms. At TCP 10 GB/s = 50 ms. Which matters for your SLA?
5. MoE expert routing changes KV access patterns. How does disaggregation behave with MoE that activates different experts per token?
## Key Terms
| Term | What people say | What it actually means |
|------|----------------|------------------------|
| Disaggregated serving | "split prefill/decode" | Separate GPU pools for each phase |
| NIXL | "NVIDIA transport" | Dynamo's inter-node KV transfer (RDMA/TCP) |
| NVIDIA Dynamo | "the orchestrator" | Stack-above coordinator for vLLM/SGLang/TRT-LLM |
| llm-d | "Kubernetes native" | Red Hat + AWS K8s disaggregated stack |
| Planner Profiler | "Dynamo auto-config" | Measures workload, configures pool ratios |
| SLA Planner | "Dynamo policy" | Auto-rate-matches prefill:decode to meet SLOs |
| `packDomain: rack` | "llm-d topology" | Pack prefill+decode on same rack for fast KV |
| UCCL | "unified collective" | llm-d 0.5 networking layer for scale-to-zero |
| MoE expert routing | "expert per token" | DeepSeek-V3 pattern; disaggregation helps |
## Further Reading
- [NVIDIA — Introducing Dynamo](https://developer.nvidia.com/blog/introducing-nvidia-dynamo-a-low-latency-distributed-inference-framework-for-scaling-reasoning-ai-models/)
- [NVIDIA — Disaggregated LLM Inference on Kubernetes](https://developer.nvidia.com/blog/deploying-disaggregated-llm-inference-workloads-on-kubernetes/)
- [TensorRT-LLM Disaggregated Serving blog](https://nvidia.github.io/TensorRT-LLM/blogs/tech_blog/blog5_Disaggregated_Serving_in_TensorRT-LLM.html)
- [llm-d GitHub](https://github.com/llm-d/llm-d)
- [llm-d 0.5 release notes](https://github.com/llm-d/llm-d/releases)
@@ -0,0 +1,31 @@
---
name: disaggregation-decider
description: Decide whether to adopt disaggregated prefill/decode (Dynamo or llm-d) for a given workload and cluster. Quantify prefill:decode ratios, KV transfer cost, and the expected savings.
version: 1.0.0
phase: 17
lesson: 17
tags: [disaggregated-serving, dynamo, llm-d, nixl, kv-transfer, prefill-decode]
---
Given workload profile (prompt/output length distribution, model, concurrency), cluster topology (GPUs, fabric, RDMA availability), and current serving cost, produce a disaggregation decision.
Produce:
1. Disaggregate? Yes / No with numbered justification. Baseline: prompts > 512 AND outputs > 200. Fabric: RDMA available helps; TCP-only pushes break-even longer.
2. Stack choice. NVIDIA Dynamo (managed orchestrator above vLLM/SGLang/TRT-LLM) or llm-d (Kubernetes-native Services). Match to the operational context.
3. Prefill:decode ratio. Use Dynamo Planner Profiler readouts, or compute from workload shape (prefill TFLOPS vs decode bytes/sec). Example: 2 prefill : 1 decode for RAG-heavy; 1:2 for output-heavy.
4. KV transfer plan. Named transport (NIXL over InfiniBand / RDMA / TCP fallback). Compute the per-request transfer tax for your prompt P99.
5. Router integration. Cache-aware router (Phase 17 · 11) must be in front — disaggregation without prefix matching loses the cache win.
6. Expected savings. Compute vs colocated baseline; cite the published case (30-40% at same SLA).
Hard rejects:
- Disaggregating short-prompt workloads (<512 tokens). Refuse — the transfer tax dominates.
- Deploying without a cache-aware router. Refuse — blind routing negates the KV locality.
- Ignoring topology (rack packing). Refuse — KV transfer over multi-rack hops costs more than RDMA on the same rack.
Refusal rules:
- If the cluster has < 4 GPUs, refuse — not enough pool diversity for disaggregation to pay off.
- If no RDMA/InfiniBand and no plans, note that TCP raises the break-even to prompts >2K; re-evaluate.
- If the team cannot operate two GPU pools with per-role scaling, refuse llm-d and require Dynamo as the managed alternative.
Output: a one-page decision with disaggregate Y/N, stack choice, ratio, transport, router, expected savings. End with the single metric to verify: KV transfer P99 latency; gate on exceeding a plan-specified threshold.