Updated · 1 episodes · 1 show · 1 source notes

concept Topics: Technology

Inference Decode Bandwidth

Definition

Inference decode bandwidth is the constraint created when autoregressive generation produces tokens sequentially and repeatedly moves model weights and state between memory and compute, making usable memory bandwidth and communication efficiency more decisive than peak arithmetic throughput.

Current Synthesis

E251 separates compute-rich prefill from bandwidth-sensitive decode. Batching lets general accelerators reuse a weight read across many requests, but it trades individual latency for utilization. SRAM-heavy systems instead bring fixed weights closer to compute, yet their limited capacity forces multi-chip topology, partitioning, and often a separate path for attention and dynamic KV cache. The relevant unit of judgment is therefore the delivered token at system level, not any one chip specification.

Key Claims

  • Autoregressive dependency limits token-level parallelism and can make decode memory-bound even when abundant compute is available.
  • Batching amortizes weight movement across requests but can increase queueing and tail latency.
  • SRAM improves locality and bandwidth but sacrifices capacity, creating multi-chip communication and model-partitioning costs.
  • KV cache and fixed model weights have different growth behavior, so heterogeneous memory and compute placement can be rational.
  • Dynamic MoE routing makes static communication schedules less efficient because expert selection is known only at runtime.
  • Token cost, latency, and power must be evaluated for the complete serving system rather than inferred from FLOPS or memory bandwidth alone.

Evidence

  • Decode mechanism and batching: E251 contrasts parallel prefill with sequential decode and explains why shared batches improve GPU utilization.
  • Memory tradeoff: E251 compares DRAM, HBM, and six-transistor SRAM across distance, bandwidth, capacity, and cost.
  • Architectural responses: E251 maps deterministic compilation to Groq, wafer-scale locality to Cerebras, and HBM plus on-chip heterogeneity to HanaPino.

Counterevidence & Qualifications

The episode’s numerical bandwidth-per-dollar gaps and model-size estimates are not third-party comparative benchmarks. Workload shape, quantization, batch size, context length, KV-cache compression, interconnect, software maturity, yield, and local electricity economics can all change the result. SRAM is therefore not a universal replacement for HBM, and decode is not the only inference phase that matters.

What Changed

  • Established a decode-specific extension of the wiki’s broader memory-wall synthesis.
  • Made batching, SRAM capacity, KV cache, MoE routing, and heterogeneous partitioning part of one system-level cost model.

Sources

1 source notes across 1 show
  1. E251|推理芯片之战:聊聊Groq、Cerebras与OpenAI三大路径与Bill Dally的设计哲学 硅谷101