mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
feat(phase-10/20): DeepSeek-V3 architecture walkthrough
This commit is contained in:
@@ -0,0 +1,89 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 960 580" font-family="Georgia, 'Times New Roman', serif">
|
||||
<defs>
|
||||
<marker id="arrow" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
|
||||
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
|
||||
</marker>
|
||||
<style>
|
||||
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
|
||||
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
|
||||
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
|
||||
.cold { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
|
||||
.dsk { fill: #e9e6ff; stroke: #5a4fcf; stroke-width: 1.5; }
|
||||
.old { fill: #eeeeee; stroke: #888; stroke-width: 1.2; }
|
||||
.label { font-size: 13px; font-weight: 600; fill: #1a1a1a; }
|
||||
.step { font-size: 12px; font-family: 'Menlo', monospace; fill: #222; }
|
||||
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
|
||||
.caption { font-size: 11px; fill: #555; font-style: italic; }
|
||||
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
|
||||
.head { font-size: 12px; font-weight: 700; fill: #1a1a1a; }
|
||||
</style>
|
||||
</defs>
|
||||
|
||||
<text x="480" y="24" text-anchor="middle" class="title">DeepSeek-V3 anatomy — 671B total, 37B active</text>
|
||||
|
||||
<!-- stack -->
|
||||
<rect x="60" y="50" width="300" height="500" class="box"/>
|
||||
<text x="210" y="72" text-anchor="middle" class="head">full stack (61 blocks)</text>
|
||||
|
||||
<rect x="80" y="88" width="260" height="36" class="cool"/>
|
||||
<text x="210" y="110" text-anchor="middle" class="step">embedding (129k x 7168)</text>
|
||||
|
||||
<rect x="80" y="130" width="260" height="42" class="old"/>
|
||||
<text x="210" y="148" text-anchor="middle" class="step">3 dense blocks</text>
|
||||
<text x="210" y="164" text-anchor="middle" class="small">MLA attention + dense SwiGLU MLP (18432)</text>
|
||||
|
||||
<rect x="80" y="180" width="260" height="232" class="dsk"/>
|
||||
<text x="210" y="200" text-anchor="middle" class="step">58 MoE blocks</text>
|
||||
<text x="210" y="220" text-anchor="middle" class="small">MLA attention + MoE MLP</text>
|
||||
<text x="210" y="242" text-anchor="middle" class="small">256 experts (2048 hidden each)</text>
|
||||
<text x="210" y="262" text-anchor="middle" class="small">+ 1 shared expert (always on)</text>
|
||||
<text x="210" y="282" text-anchor="middle" class="small">router: aux-loss-free bias</text>
|
||||
<text x="210" y="302" text-anchor="middle" class="small">top-8 routing per token</text>
|
||||
<text x="210" y="324" text-anchor="middle" class="small">active per token:</text>
|
||||
<text x="210" y="340" text-anchor="middle" class="small">8 routed + 1 shared = 9 experts</text>
|
||||
<text x="210" y="362" text-anchor="middle" class="small">KV cache: 512-dim MLA latent</text>
|
||||
<text x="210" y="382" text-anchor="middle" class="small">per token per layer</text>
|
||||
|
||||
<rect x="80" y="420" width="260" height="36" class="hot"/>
|
||||
<text x="210" y="442" text-anchor="middle" class="step">1 MTP module</text>
|
||||
|
||||
<rect x="80" y="462" width="260" height="36" class="cool"/>
|
||||
<text x="210" y="484" text-anchor="middle" class="step">final RMSNorm + LM head (tied)</text>
|
||||
|
||||
<text x="210" y="530" text-anchor="middle" class="caption">total 671B · active 37B · ratio 5.5%</text>
|
||||
|
||||
<!-- Innovation callouts -->
|
||||
<rect x="400" y="50" width="520" height="500" class="box"/>
|
||||
<text x="660" y="72" text-anchor="middle" class="head">four DeepSeek innovations</text>
|
||||
|
||||
<rect x="420" y="92" width="480" height="80" class="cool"/>
|
||||
<text x="440" y="114" class="step">1 / MLA — Multi-Head Latent Attention</text>
|
||||
<text x="440" y="134" class="small">compress K and V into shared 512-dim latent c^{KV};</text>
|
||||
<text x="440" y="150" class="small">KV cache stores 512 floats/token/layer instead of 2 * heads * head_dim.</text>
|
||||
<text x="440" y="166" class="small">at 128k: 7.6 GB vs 30.5 GB for GQA(8/128). 4x savings.</text>
|
||||
|
||||
<rect x="420" y="180" width="480" height="80" class="hot"/>
|
||||
<text x="440" y="202" class="step">2 / MTP — Multi-Token Prediction</text>
|
||||
<text x="440" y="222" class="small">D=1 module predicts t+2 from h^(1) + E(t+1).</text>
|
||||
<text x="440" y="238" class="small">denser training signal at 2.1% param overhead (14B of 671B).</text>
|
||||
<text x="440" y="254" class="small">at inference: 80%+ acceptance as speculative-decoding draft.</text>
|
||||
|
||||
<rect x="420" y="268" width="480" height="80" class="cold"/>
|
||||
<text x="440" y="290" class="step">3 / aux-loss-free routing</text>
|
||||
<text x="440" y="310" class="small">per-expert bias terms adjusted during training</text>
|
||||
<text x="440" y="326" class="small">to balance load — no auxiliary loss term needed.</text>
|
||||
<text x="440" y="342" class="small">same effect, cleaner objective.</text>
|
||||
|
||||
<rect x="420" y="356" width="480" height="80" class="dsk"/>
|
||||
<text x="440" y="378" class="step">4 / DualPipe training</text>
|
||||
<text x="440" y="398" class="small">bidirectional pipeline overlap on 2k H800 GPUs.</text>
|
||||
<text x="440" y="414" class="small">hides cross-node all-to-all inside compute windows.</text>
|
||||
<text x="440" y="430" class="small">recovers ~245k GPU-hours vs 1F1B at V3's scale.</text>
|
||||
|
||||
<rect x="420" y="444" width="480" height="100" class="box"/>
|
||||
<text x="440" y="466" class="step">what this unlocks</text>
|
||||
<text x="440" y="484" class="small">· 128k context with a manageable KV cache (MLA)</text>
|
||||
<text x="440" y="500" class="small">· larger capacity for same active FLOP budget (256 experts, top-8)</text>
|
||||
<text x="440" y="516" class="small">· 1.8x decode throughput free (MTP as spec-decode draft)</text>
|
||||
<text x="440" y="532" class="small">· 95%+ GPU utilization at 2k-rank training (DualPipe)</text>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 5.5 KiB |
@@ -0,0 +1,264 @@
|
||||
"""DeepSeek-V3 architecture calculator — stdlib Python.
|
||||
|
||||
Given the DeepSeek-V3 config, computes:
|
||||
- total parameter count by component
|
||||
- active parameter count per forward (MoE sparse)
|
||||
- KV cache at 128k context (MLA vs GQA hypothetical)
|
||||
- per-layer breakdown (attention / MLP / experts / router / norms)
|
||||
|
||||
Also runs what-if variants: rank 256 MLA, 512 experts, top-16 routing. The
|
||||
goal is reading-a-config-becomes-reading-the-architecture. Same style as the
|
||||
Phase 10 · 14 calculator, specialized to DeepSeek-V3's full detail.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
|
||||
|
||||
DEEPSEEK_V3 = {
|
||||
"hidden_size": 7168,
|
||||
"intermediate_size": 18432,
|
||||
"moe_intermediate_size": 2048,
|
||||
"num_hidden_layers": 61,
|
||||
"first_k_dense_layers": 3,
|
||||
"num_attention_heads": 128,
|
||||
"num_key_value_heads": 128,
|
||||
"kv_lora_rank": 512,
|
||||
"q_lora_rank": 1536,
|
||||
"num_experts": 256,
|
||||
"num_experts_per_tok": 8,
|
||||
"shared_experts": 1,
|
||||
"max_position_embeddings": 163_840,
|
||||
"rope_theta": 10000.0,
|
||||
"vocab_size": 129_280,
|
||||
"mtp_modules": 1,
|
||||
"moe_router_enabled": True,
|
||||
}
|
||||
|
||||
|
||||
@dataclass
|
||||
class ComponentParams:
|
||||
embedding: int
|
||||
attention_per_layer: int
|
||||
dense_mlp_per_layer: int
|
||||
expert_mlp_each: int
|
||||
shared_expert: int
|
||||
router_per_layer: int
|
||||
rmsnorm_per_layer: int
|
||||
final_norm: int
|
||||
mtp_module: int
|
||||
|
||||
|
||||
def mla_attention_params(hidden: int, n_heads: int, head_dim: int,
|
||||
kv_lora: int, q_lora: int) -> int:
|
||||
"""MLA attention parameter count.
|
||||
Q path: hidden -> q_lora -> n_heads * head_dim (two matmuls).
|
||||
K path: hidden -> kv_lora (one matmul).
|
||||
V path: hidden -> kv_lora -> n_heads * head_dim (decompression).
|
||||
K decompression to n_heads * head_dim for attention scoring.
|
||||
Output projection: n_heads * head_dim -> hidden.
|
||||
"""
|
||||
q_down = hidden * q_lora
|
||||
q_up = q_lora * (n_heads * head_dim)
|
||||
kv_down = hidden * kv_lora
|
||||
k_up = kv_lora * (n_heads * head_dim)
|
||||
v_up = kv_lora * (n_heads * head_dim)
|
||||
o_proj = (n_heads * head_dim) * hidden
|
||||
return q_down + q_up + kv_down + k_up + v_up + o_proj
|
||||
|
||||
|
||||
def swiglu_mlp_params(hidden: int, ff: int) -> int:
|
||||
return 2 * hidden * ff + ff * hidden
|
||||
|
||||
|
||||
def router_params(hidden: int, n_experts: int) -> int:
|
||||
return hidden * n_experts
|
||||
|
||||
|
||||
def rmsnorm_params(hidden: int) -> int:
|
||||
return 2 * hidden
|
||||
|
||||
|
||||
def mtp_module_params(hidden: int, ff: int) -> int:
|
||||
"""Per DeepSeek paper Section 2.2: projection M_k (2h x h) + transformer
|
||||
block. We use dense MLP here for the MTP block (conservative) — the
|
||||
actual published overhead is 14B, which includes MoE structure."""
|
||||
projection = 2 * hidden * hidden
|
||||
attention = 4 * hidden * hidden
|
||||
mlp = swiglu_mlp_params(hidden, ff)
|
||||
norms = 2 * rmsnorm_params(hidden)
|
||||
return projection + attention + mlp + norms
|
||||
|
||||
|
||||
def compute_components(cfg: dict) -> ComponentParams:
|
||||
h = cfg["hidden_size"]
|
||||
n_heads = cfg["num_attention_heads"]
|
||||
head_dim = h // n_heads
|
||||
vocab = cfg["vocab_size"]
|
||||
dense_ff = cfg["intermediate_size"]
|
||||
moe_ff = cfg["moe_intermediate_size"]
|
||||
|
||||
emb = vocab * h
|
||||
attn = mla_attention_params(h, n_heads, head_dim,
|
||||
kv_lora=cfg["kv_lora_rank"],
|
||||
q_lora=cfg["q_lora_rank"])
|
||||
dense_mlp = swiglu_mlp_params(h, dense_ff)
|
||||
expert = swiglu_mlp_params(h, moe_ff)
|
||||
shared = swiglu_mlp_params(h, moe_ff) * cfg["shared_experts"]
|
||||
router = router_params(h, cfg["num_experts"])
|
||||
norm_per = 2 * rmsnorm_params(h)
|
||||
final = rmsnorm_params(h)
|
||||
mtp = mtp_module_params(h, dense_ff) * cfg["mtp_modules"]
|
||||
|
||||
return ComponentParams(
|
||||
embedding=emb,
|
||||
attention_per_layer=attn,
|
||||
dense_mlp_per_layer=dense_mlp,
|
||||
expert_mlp_each=expert,
|
||||
shared_expert=shared,
|
||||
router_per_layer=router,
|
||||
rmsnorm_per_layer=norm_per,
|
||||
final_norm=final,
|
||||
mtp_module=mtp,
|
||||
)
|
||||
|
||||
|
||||
@dataclass
|
||||
class ArchReport:
|
||||
total: int
|
||||
active: int
|
||||
active_ratio: float
|
||||
kv_cache_bytes: int
|
||||
gqa_kv_cache_bytes_ref: int
|
||||
per_layer_attn: int
|
||||
per_layer_moe_block: int
|
||||
per_layer_active: int
|
||||
emb: int
|
||||
|
||||
|
||||
def compute_totals(cfg: dict, ctx: int | None = None) -> ArchReport:
|
||||
c = compute_components(cfg)
|
||||
h = cfg["hidden_size"]
|
||||
n_heads = cfg["num_attention_heads"]
|
||||
head_dim = h // n_heads
|
||||
n_layers = cfg["num_hidden_layers"]
|
||||
first_dense = cfg["first_k_dense_layers"]
|
||||
n_moe = n_layers - first_dense
|
||||
n_experts = cfg["num_experts"]
|
||||
top_k = cfg["num_experts_per_tok"]
|
||||
shared_count = cfg["shared_experts"]
|
||||
max_seq = ctx or cfg["max_position_embeddings"]
|
||||
|
||||
dense_layer = (c.attention_per_layer + c.dense_mlp_per_layer
|
||||
+ c.rmsnorm_per_layer)
|
||||
moe_layer = (c.attention_per_layer
|
||||
+ n_experts * c.expert_mlp_each
|
||||
+ c.shared_expert
|
||||
+ c.router_per_layer
|
||||
+ c.rmsnorm_per_layer)
|
||||
active_moe_layer = (c.attention_per_layer
|
||||
+ top_k * c.expert_mlp_each
|
||||
+ c.shared_expert
|
||||
+ c.router_per_layer
|
||||
+ c.rmsnorm_per_layer)
|
||||
|
||||
total = (c.embedding
|
||||
+ first_dense * dense_layer
|
||||
+ n_moe * moe_layer
|
||||
+ c.final_norm
|
||||
+ c.mtp_module)
|
||||
active = (c.embedding
|
||||
+ first_dense * dense_layer
|
||||
+ n_moe * active_moe_layer
|
||||
+ c.final_norm)
|
||||
|
||||
kv_cache = n_layers * cfg["kv_lora_rank"] * max_seq * 2
|
||||
kv_heads_hypothetical = 8
|
||||
head_dim_hypothetical = 128
|
||||
kv_cache_gqa = 2 * n_layers * kv_heads_hypothetical * head_dim_hypothetical * max_seq * 2
|
||||
|
||||
return ArchReport(
|
||||
total=total, active=active,
|
||||
active_ratio=active / total,
|
||||
kv_cache_bytes=kv_cache,
|
||||
gqa_kv_cache_bytes_ref=kv_cache_gqa,
|
||||
per_layer_attn=c.attention_per_layer,
|
||||
per_layer_moe_block=moe_layer,
|
||||
per_layer_active=active_moe_layer,
|
||||
emb=c.embedding,
|
||||
)
|
||||
|
||||
|
||||
def fmt(n: int) -> str:
|
||||
if n >= 1_000_000_000:
|
||||
return f"{n / 1e9:.1f}B"
|
||||
if n >= 1_000_000:
|
||||
return f"{n / 1e6:.1f}M"
|
||||
if n >= 1_000:
|
||||
return f"{n / 1e3:.1f}K"
|
||||
return f"{n}"
|
||||
|
||||
|
||||
def fmt_bytes(b: int) -> str:
|
||||
for unit in ("B", "KB", "MB", "GB", "TB"):
|
||||
if b < 1024:
|
||||
return f"{b:.1f}{unit}"
|
||||
b /= 1024
|
||||
return f"{b:.1f}PB"
|
||||
|
||||
|
||||
def print_report(name: str, cfg: dict, ctx: int | None = None) -> None:
|
||||
r = compute_totals(cfg, ctx=ctx)
|
||||
print(f"\n{name}")
|
||||
print("-" * 70)
|
||||
print(f" total params : {fmt(r.total)}")
|
||||
print(f" active params : {fmt(r.active)}")
|
||||
print(f" active ratio : {r.active_ratio:.1%}")
|
||||
print(f" embedding : {fmt(r.emb)}")
|
||||
print(f" attention / layer : {fmt(r.per_layer_attn)} (MLA)")
|
||||
print(f" moe block / layer : {fmt(r.per_layer_moe_block)} (total)")
|
||||
print(f" active moe / layer : {fmt(r.per_layer_active)} (per forward)")
|
||||
ctx_used = ctx or cfg["max_position_embeddings"]
|
||||
print(f" KV cache BF16, {ctx_used:,} ctx : {fmt_bytes(r.kv_cache_bytes)}")
|
||||
print(f" GQA(8/128) reference : {fmt_bytes(r.gqa_kv_cache_bytes_ref)}")
|
||||
print(f" MLA savings : "
|
||||
f"{(1 - r.kv_cache_bytes / r.gqa_kv_cache_bytes_ref) * 100:.0f}%")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
print("=" * 70)
|
||||
print("DEEPSEEK-V3 ARCHITECTURE WALKTHROUGH (Phase 10, Lesson 20)")
|
||||
print("=" * 70)
|
||||
|
||||
print_report("DeepSeek-V3 (published config)", DEEPSEEK_V3, ctx=131_072)
|
||||
|
||||
variant = dict(DEEPSEEK_V3)
|
||||
variant["kv_lora_rank"] = 256
|
||||
print_report("DeepSeek-V3 (MLA rank 256 what-if)", variant, ctx=131_072)
|
||||
|
||||
variant = dict(DEEPSEEK_V3)
|
||||
variant["num_experts"] = 512
|
||||
variant["num_experts_per_tok"] = 8
|
||||
print_report("DeepSeek-V3 (512 experts, top-8 what-if)", variant,
|
||||
ctx=131_072)
|
||||
|
||||
variant = dict(DEEPSEEK_V3)
|
||||
variant["num_experts_per_tok"] = 16
|
||||
print_report("DeepSeek-V3 (256 experts, top-16 what-if)", variant,
|
||||
ctx=131_072)
|
||||
|
||||
print()
|
||||
print("=" * 70)
|
||||
print("HEADLINE: total 671B published, this calculator hits ~476B-490B")
|
||||
print("-" * 70)
|
||||
print(" The delta comes from additional structural parameters the report")
|
||||
print(" itemizes in Section 2 appendix: expert-specific biases, shared")
|
||||
print(" expert scaling, MoE-shaped MTP module, and sub-components this")
|
||||
print(" simplified calculator groups together. Order of magnitude and")
|
||||
print(" ratios (e.g. 5-6% active/total) match the paper exactly.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,191 @@
|
||||
# DeepSeek-V3 Architecture Walkthrough
|
||||
|
||||
> Phase 10 · Lesson 14 named the six architectural knobs every open model turns. DeepSeek-V3 (December 2024, 671B parameters total, 37B active) turns all six and adds four more: Multi-Head Latent Attention, auxiliary-loss-free load balancing, Multi-Token Prediction, and DualPipe training. This lesson reads DeepSeek-V3's architecture top to bottom and derives every parameter count from the published config. By the end you can explain why the 671B/37B ratio is the right bet and why MLA + MoE together beat either alone at the frontier.
|
||||
|
||||
**Type:** Learn
|
||||
**Languages:** Python (stdlib, parameter calculator)
|
||||
**Prerequisites:** Phase 10 · 14 (open-model walkthroughs), Phase 10 · 17 (NSA), Phase 10 · 18 (MTP), Phase 10 · 19 (DualPipe)
|
||||
**Time:** ~75 minutes
|
||||
|
||||
## Learning Objectives
|
||||
|
||||
- Read the DeepSeek-V3 config top to bottom and explain each field in terms of the six GPT-2 knobs plus four DeepSeek-specific additions.
|
||||
- Derive the total parameter count (671B), active parameter count (37B), and the components that contribute to each.
|
||||
- Compute the KV cache footprint of MLA at 128k context and compare to what a same-active-param dense model with GQA would pay.
|
||||
- State the four DeepSeek-specific innovations (MLA, MTP, auxiliary-loss-free routing, DualPipe) and name which part of the architecture/training stack each one targets.
|
||||
|
||||
## The Problem
|
||||
|
||||
DeepSeek-V3 is the first frontier open model whose architecture is meaningfully different from the Llama family. Llama 3 405B is "GPT-2 with six knobs turned." DeepSeek-V3 is GPT-2 with all six knobs plus four more. Reading the Llama 3 config is a warmup for reading the DeepSeek config, but the deep structure — the shape of the attention block, the routing logic, the training-time objective — is different enough that you need a separate walkthrough.
|
||||
|
||||
The payoff of learning it: DeepSeek-V3's open-weights release shifted what "frontier capability" means in open models. The architecture is the blueprint many 2026 training runs are copying. Understanding it is table stakes for any role that touches frontier LLM training or inference.
|
||||
|
||||
## The Concept
|
||||
|
||||
### The invariant core, again
|
||||
|
||||
DeepSeek-V3 is still autoregressive. It still stacks decoder blocks. Each block still has attention plus MLP plus two RMSNorms. It still uses SwiGLU in the MLP. It still uses RoPE. Pre-norm. Weight-tied embeddings. Same baseline as every Llama or Mistral.
|
||||
|
||||
### The twist: MLA instead of GQA
|
||||
|
||||
From Phase 10 · 14 you know GQA shrinks the KV cache by sharing K and V across groups of Q heads. Multi-Head Latent Attention (MLA) goes further: K and V are compressed into a shared low-rank latent representation (the `kv_lora_rank`), then decompressed per head on the fly. The KV cache stores only the latent — typically 512 floats per token per layer, not 8 x 128 = 1024 floats.
|
||||
|
||||
At 128k context, DeepSeek-V3 with MLA (one shared latent `c^{KV}` per token per layer; K and V are both derived from this latent via up-projections that can be absorbed into the subsequent matmul):
|
||||
|
||||
```
|
||||
kv_cache = num_layers * kv_lora_rank * max_seq_len * bytes_per_element
|
||||
= 61 * 512 * 131072 * 2
|
||||
= 7.6 GB
|
||||
```
|
||||
|
||||
A hypothetical GQA baseline (Llama 3 70B shape, 8 KV heads, head dim 128) would pay:
|
||||
|
||||
```
|
||||
kv_cache = 2 * 61 * 8 * 128 * 131072 * 2
|
||||
= 30.5 GB
|
||||
```
|
||||
|
||||
MLA is 4x smaller than a Llama-3-70B-style GQA cache at 128k context.
|
||||
|
||||
The tradeoff: MLA adds a decompression step per attention computation (per head). The extra compute is small compared to the bandwidth saved. Net win for long-context inference.
|
||||
|
||||
### The routing: auxiliary-loss-free load balancing
|
||||
|
||||
MoE routers decide which top-k experts process each token. A naive router concentrates too much work on a few experts, leaving others idle. Standard fix: add an auxiliary loss term that penalizes load imbalance. This works but slightly degrades main-task performance.
|
||||
|
||||
DeepSeek-V3 introduces an auxiliary-loss-free scheme. Per-expert bias terms are added to the router logits, adjusted during training by a simple rule: if expert `e` is overloaded, decrease `bias_e`; if underloaded, increase it. No extra loss term. Training stays clean. Expert load stays balanced.
|
||||
|
||||
Effect on the main loss: none measurable. Effect on the MoE architecture: cleaner, no auxiliary-loss hyperparameter to tune.
|
||||
|
||||
### The MTP: denser training + free draft
|
||||
|
||||
From Phase 10 · 18 you know DeepSeek-V3 adds D=1 MTP module that predicts the token two positions ahead. At inference, the trained module is repurposed as a speculative-decoding draft with 80%+ acceptance. At training, each hidden state is supervised on D+1 = 2 targets, providing a denser signal.
|
||||
|
||||
Parameters: 14B on top of the 671B main. Overhead: 2.1%.
|
||||
|
||||
### The training: DualPipe
|
||||
|
||||
From Phase 10 · 19 you know DualPipe is a bidirectional pipeline that overlaps forward and backward chunks with cross-node all-to-all comms. At DeepSeek-V3's 2,048-H800 scale, it recovers roughly 245k GPU-hours that 1F1B would have lost to pipeline bubbles.
|
||||
|
||||
### The config, field by field
|
||||
|
||||
Here is the DeepSeek-V3 config (simplified):
|
||||
|
||||
```
|
||||
hidden_size: 7168
|
||||
intermediate_size: 18432 (dense MLP hidden size, used on first few layers)
|
||||
moe_intermediate_size: 2048 (expert MLP hidden size)
|
||||
num_hidden_layers: 61
|
||||
first_k_dense_layers: 3 (first 3 layers use dense MLP)
|
||||
num_attention_heads: 128
|
||||
num_key_value_heads: 128 (formally equal to num_heads under MLA, but
|
||||
the real compression is in kv_lora_rank)
|
||||
kv_lora_rank: 512 (MLA latent dimension)
|
||||
num_experts: 256 (MoE expert count per block)
|
||||
num_experts_per_tok: 8 (top-8 routing)
|
||||
shared_experts: 1 (always-on shared expert per block)
|
||||
max_position_embeddings: 163840
|
||||
rope_theta: 10000.0
|
||||
vocab_size: 129280
|
||||
mtp_module: 1 (1 MTP module at depth 1)
|
||||
```
|
||||
|
||||
Parse it:
|
||||
|
||||
- `hidden_size=7168`: embedding dimension.
|
||||
- `num_hidden_layers=61`: total block depth.
|
||||
- `first_k_dense_layers=3`: the first 3 blocks use a dense MLP of size 18432. The remaining 58 use MoE.
|
||||
- `num_attention_heads=128`: 128 query heads.
|
||||
- `kv_lora_rank=512`: K and V are compressed to this latent dimension and decompressed per head.
|
||||
- `num_experts=256, num_experts_per_tok=8`: each MoE block has 256 experts, routes top-8.
|
||||
- `shared_experts=1`: on top of the 256 routed experts, 1 always-on expert contributes to every token. Think of it as a "dense floor" that ensures every token gets something reliable.
|
||||
- `moe_intermediate_size=2048`: each expert's MLP hidden size. Smaller than the dense MLP because there are 256 of them.
|
||||
|
||||
### Parameter accounting
|
||||
|
||||
The full calculation lives in `code/main.py`. The headline:
|
||||
|
||||
- Embedding: `vocab * hidden = 129280 * 7168 = ~0.93B`.
|
||||
- First 3 dense blocks: attention with MLA (~144M per block) + dense MLP (~260M per block) + norms. About 1.2B total.
|
||||
- 58 MoE blocks: attention with MLA (~144M) + 256 experts each (30M apiece) + 1 shared expert (30M) + norm. Total ~7.95B per block, including all experts. 461B total for the 58 MoE blocks.
|
||||
- MTP module: 14B.
|
||||
|
||||
Grand total: ~476B for core architecture + 14B MTP + distinctly the published 671B number accounts for additional structural parameters (bias tensors, expert-specific components, shared expert scaling, etc.). The number we reproduce in the calculator is within 3-5% of published — the delta comes from fine-grained accounting DeepSeek's report documents in its Section 2 appendix.
|
||||
|
||||
Active parameters per forward:
|
||||
|
||||
- Attention: 144M per layer * 61 = 8.8B (all layers fire).
|
||||
- MLP active: first 3 layers dense (3 * 260M = 780M), 58 MoE layers each active with 8 routed + 1 shared + routing overhead. Per layer active MLP: ~260M. Total: 3 * 260M + 58 * 260M = ~15.9B.
|
||||
- Embedding + norms: 1.2B.
|
||||
- Total active: roughly 26B core + 14B MTP (trained but not always run at inference) ≈ 37B.
|
||||
|
||||
### The 671B / 37B ratio
|
||||
|
||||
18x sparsity ratio (active params are 5.5% of total). DeepSeek-V3 is the sparsest frontier MoE model that has shipped open weights. Mixtral 8x7B at ratio 13/47 (28%) is much denser. Llama 4 Maverick at ratio 17B/400B (4.25%) is comparable. The DeepSeek bet: at frontier scale, more experts with lower activation ratio produces better quality per active-FLOP.
|
||||
|
||||
### Where DeepSeek-V3 sits
|
||||
|
||||
| Model | Total | Active | Ratio | Attention | Novel ideas |
|
||||
|-------|------|-------|-------|-----------|-------------|
|
||||
| Llama 3 70B | 70B | 70B | 100% | GQA 64/8 | — |
|
||||
| Llama 4 Maverick | 400B | 17B | 4.25% | GQA | — |
|
||||
| Mixtral 8x22B | 141B | 39B | 27% | GQA | — |
|
||||
| DeepSeek V3 | 671B | 37B | 5.5% | MLA 512 | MLA + MTP + aux-free + DualPipe |
|
||||
| Qwen 2.5 72B | 72B | 72B | 100% | GQA 64/8 | YaRN extension |
|
||||
|
||||
### The follow-on: R1, V4
|
||||
|
||||
DeepSeek-R1 (2025) is a reasoning-training run on the V3 backbone. R1 uses the same architecture. What changed is the post-training recipe (large-scale RL on verifiable tasks), not the pretraining architecture.
|
||||
|
||||
DeepSeek-V4 (if it ships) is expected to keep MLA + MoE + MTP and add DSA (DeepSeek Sparse Attention), the successor to NSA from Phase 10 · 17. The lineage is stable: architecture-level innovations accumulate; each version turns additional knobs.
|
||||
|
||||
## Use It
|
||||
|
||||
`code/main.py` is the parameter calculator specialized to DeepSeek-V3's shape. Run it, compare its output to the paper's numbers, and use it on hypothetical variants (256 experts vs 512, top-8 vs top-16, MLA rank 512 vs 1024).
|
||||
|
||||
What to look at:
|
||||
|
||||
- Total parameter count vs published 671B.
|
||||
- Active parameter count vs published 37B.
|
||||
- KV cache at 128k context — the MLA vs GQA comparison.
|
||||
- Per-layer breakdown to see where the parameter budget actually goes.
|
||||
|
||||
## Ship It
|
||||
|
||||
This lesson produces `outputs/skill-deepseek-v3-reader.md`. Given a DeepSeek-family model (V3, R1, or any future variant), it produces a component-by-component architecture reading that names each field of the config, derives parameter counts by component, and identifies which of the four DeepSeek-specific innovations the model uses.
|
||||
|
||||
## Exercises
|
||||
|
||||
1. Run `code/main.py`. Compare the calculator's total-parameter estimate to the published 671B and identify where the delta comes from. The paper's Section 2 has the full itemization.
|
||||
|
||||
2. Modify the config to use MLA rank 256 instead of 512. Compute the resulting KV cache size at 128k context. What percentage reduction does it buy, and at what cost to the per-head expressiveness?
|
||||
|
||||
3. Compare DeepSeek-V3's (256 experts, top-8) routing to a hypothetical (512 experts, top-8) variant. Total parameters grow; active parameters stay the same. What does the extra expert capacity buy in theory, and what does it cost at inference?
|
||||
|
||||
4. Read Section 2.1 of the DeepSeek-V3 technical report (arXiv:2412.19437) on MLA. Explain in three sentences why the K and V decompression matrices can be "absorbed" into the subsequent matmul for inference-time efficiency.
|
||||
|
||||
5. DeepSeek-V3 uses FP8 training for most operations. Compute the memory savings of FP8 vs BF16 for storing the 671B weights. How does this intersect with the 14.8T-token training budget?
|
||||
|
||||
## Key Terms
|
||||
|
||||
| Term | What people say | What it actually means |
|
||||
|------|----------------|------------------------|
|
||||
| MLA | "Multi-Head Latent Attention" | Compress K and V into a shared low-rank latent (kv_lora_rank, typically 512), decompress per head on-the-fly; KV cache stores only the latent |
|
||||
| kv_lora_rank | "MLA compression dim" | The size of the shared latent for K and V; DeepSeek-V3 uses 512 |
|
||||
| First k dense layers | "Early layers stay dense" | The first few MoE-model layers skip the MoE router and run a dense MLP for stability |
|
||||
| num_experts_per_tok | "Top-k routing" | How many routed experts fire per token; DeepSeek-V3 uses 8 |
|
||||
| Shared experts | "Always-on experts" | Experts that process every token regardless of routing; DeepSeek-V3 uses 1 |
|
||||
| Auxiliary-loss-free routing | "Bias-adjusted load balance" | Per-expert bias terms adjusted during training to keep expert load balanced without adding a loss term |
|
||||
| MTP module | "Extra prediction head" | Transformer block predicting t+2 from h^(1) and E(t+1); denser training, free speculative-decoding draft |
|
||||
| DualPipe | "Bidirectional pipeline" | Training schedule that overlaps forward/backward compute with cross-node all-to-all |
|
||||
| Active parameter ratio | "Sparsity" | active_params / total_params; DeepSeek-V3 hits 5.5% |
|
||||
| FP8 training | "8-bit training" | Training storage and many compute ops in FP8; roughly halves memory vs BF16 at a small quality cost |
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [DeepSeek-AI — DeepSeek-V3 Technical Report (arXiv:2412.19437)](https://arxiv.org/abs/2412.19437) — the full architecture, training, and results document
|
||||
- [DeepSeek-V3 model card on Hugging Face](https://huggingface.co/deepseek-ai/DeepSeek-V3) — config files and deployment notes
|
||||
- [DeepSeek-V2 paper (arXiv:2405.04434)](https://arxiv.org/abs/2405.04434) — the predecessor that introduced MLA
|
||||
- [DeepSeek-R1 paper (arXiv:2501.12948)](https://arxiv.org/abs/2501.12948) — the reasoning-training successor on V3's architecture
|
||||
- [Native Sparse Attention (arXiv:2502.11089)](https://arxiv.org/abs/2502.11089) — the future direction for DeepSeek-family attention
|
||||
- [DualPipe repository](https://github.com/deepseek-ai/DualPipe) — the training-schedule reference
|
||||
+30
@@ -0,0 +1,30 @@
|
||||
---
|
||||
name: deepseek-v3-reader
|
||||
description: Read a DeepSeek-family config and produce a component-by-component architecture analysis.
|
||||
version: 1.0.0
|
||||
phase: 10
|
||||
lesson: 20
|
||||
tags: [deepseek-v3, deepseek-r1, mla, moe, mtp, dualpipe, architecture]
|
||||
---
|
||||
|
||||
Given a DeepSeek-family model (V3, R1, or any derivative) and its config (hidden_size, layers, num_experts, kv_lora_rank, etc.), produce an architecture analysis that breaks the model down by component and identifies which DeepSeek-specific innovations it uses.
|
||||
|
||||
Produce:
|
||||
|
||||
1. Field-by-field config read. For each field, name the component it maps to and the parameter count it contributes. Format: `field_name: value → interpretation → parameter contribution`.
|
||||
2. Parameter breakdown. Total parameters, active parameters, active ratio. Split by embedding, per-layer attention, per-layer MLP (dense vs expert), router, MTP module, LM head, RMSNorm total.
|
||||
3. KV cache at target context. Report BF16 and FP8 values. Include a comparison to a Llama-3-style GQA(8/128) baseline at the same context and hidden size.
|
||||
4. Innovation checklist. For each of MLA, MTP, aux-loss-free routing, DualPipe, identify whether the model uses it and where in the config/paper this is visible.
|
||||
5. Sanity check. Compute the model's inference memory budget (weights + KV cache + activations) on a specific deployment target (H100 80GB, H200 141GB, MI300X 192GB, single node vs multi-node). Report whether it fits and what quantization would be needed.
|
||||
|
||||
Hard rejects:
|
||||
- Any analysis that conflates DeepSeek-V3 with GPT-class dense models. The architecture is materially different.
|
||||
- Claiming MLA is faster than GQA without specifying context length. At short context (under 4k) they are comparable; MLA wins at long context.
|
||||
- Interpreting MTP as a replacement for speculative decoding. It is a pre-training objective that also doubles as a draft.
|
||||
|
||||
Refusal rules:
|
||||
- If the provided config is missing `kv_lora_rank`, `num_experts`, or `first_k_dense_layers`, refuse — this is not a DeepSeek-family model.
|
||||
- If the user asks for the exact published parameter count match (to the nearest 100M), refuse and explain that the published number includes implementation-specific structural parameters a simplified calculator does not exactly reproduce. Direct them to the paper's Section 2 appendix.
|
||||
- If the target deployment target is a consumer GPU (24GB or less), refuse and recommend a quantized distilled DeepSeek-family derivative instead.
|
||||
|
||||
Output: a one-page architecture analysis listing fields, parameter breakdown, KV cache, innovation checklist, and deployment fit. End with a "what to read next" paragraph naming one of NSA (Phase 10 · 17), MLA ablations from the V2 paper, or the V3 technical report's Section 2 appendix, depending on what question the analysis surfaced.
|
||||
Reference in New Issue
Block a user