DeepSeek V4.1 Flash inference explorer · memory tiers

A deliberately small 5-layer causal encoder + 5-layer decoder schematic. It illustrates dataflow, not exact layer placement or dimensions.
Global KV 0
SWA KV 0
Engram rows 0
Token flow and state placementREADY

INPUT / OUTPUT TOKENS

CAUSAL ENCODER — runs in prefill

Encoder final hidden stream H[t]

DECODER — skipped in prompt prefill

Decoder reads projected global KV

MEMORY HIERARCHY — HARDWARE OVERLAY

width ≈ capacity · stripe ≈ latency · five orders of magnitude top to bottom
GPU SRAM~1–5 ns~50 MB · 20–100 TB/s on-chip
Kernel working set — activations, indexer scores, Top-K sort
GPU HBM3e~0.4–0.7 µs141–192 GB/GPU · 5–8 TB/s
MoE weights — 552B total · 16B active

Global KV — active sessions · FP4 · token-addressable

FP4: compact retained scalar values

CSA2 sparse index — selected global positions

SWA KV — active window
Host DRAM (PCIe 5 / CXL)~1–2 µsTBs/host · ~64 GB/s hop

SWA KV — recent-session pool (10% of host DRAM)

Engram — 196B learned weight tables (not activations)

Model weights (not activations) · hash current n-grams → sparse rows → learned gate
NVMe SSD (Gen5)~50–100 µstens of TB · ~12–14 GB/s per drive
Global KV — persistent prefix cache (≥ 72 h)
Start here. Prefill streams the prompt through the causal encoder. Click an animation, then inspect the highlighted state.