Updated · 8 episodes · 5 shows · 8 source notes

concept Topics: Technology

High Bandwidth Memory

Definition

High bandwidth memory is stacked, high-throughput memory placed close to AI accelerators so models can move parameters, activations, and inference cache data fast enough for training and inference workloads.

Current Synthesis

HBM is now one of the wiki’s clearest physical AI bottlenecks. Marketplace sources explain that AI chips need fast memory close to accelerators and that supply is concentrated among Micron Technology, SK Hynix, and Samsung. Technical sources add that HBM sits inside a wider memory hierarchy with advanced packaging, memory-wall pressure, CXL, flash, and accelerator roadmaps. The new All-In episode sharpens the market-cycle view: 8-high stacks are moving toward 12- and 16-high stacks, Micron’s 2026 supply is described as sold out, and Gavin Baker calls DRAM/HBM capacity one of AI’s most important bottlenecks.

Key Claims

  • AI acceleration depends on fast memory bandwidth and capacity as well as raw compute.
  • HBM demand can lift memory suppliers such as Micron Technology, SK Hynix, and Samsung when AI data-center buildout accelerates.
  • Supplier concentration, advanced packaging, yield, and fab capital intensity make HBM shortages hard to solve quickly.
  • HBM demand is intensified by inference, long-context KV cache, and larger model-memory footprints, not only training.
  • AI memory demand can spill into consumer markets by changing product focus, capacity allocation, and pricing.
  • Alternative memory architectures can improve utilization but do not replace HBM in the hottest low-latency layer in the source set.
  • Local fab expansion can make HBM capacity a community-benefit, water, emissions, labor, and land-use issue.

Evidence

Counterevidence & Qualifications

HBM is a bottleneck but not the only one: advanced packaging, GPUs, networking, power, cooling, site readiness, and demand timing all affect delivered AI capacity. Some source figures are product or market commentary rather than filings. Memory is cyclical, so supplier strength can reverse if supply catches demand or AI capex slows.

What Changed

  • Migrated the page to the synthesis-first concept schema.
  • Added Micron’s All-In quarter and sold-out 2026 supply as source-scoped evidence.
  • Added stacked DRAM progression from 8-high to 12- and 16-high HBM stacks.
  • Elevated inference and consumer-device spillover in the current synthesis.

Sources

8 source notes across 5 shows
  1. Bytes: Week in Review - SpaceX eyes an IPO, community members want legal commitments from Micron, and YouTube to ditch AI slop Marketplace Tech
  2. AI is eating up the world's computing memory Marketplace Tech
  3. E230|1万亿收入预期背后:英伟达的巅峰与软肋 硅谷101
  4. Bytes: Week in Review - Micron''s big earnings, Oracle''s data center woes and "slop" is Merriam-Webster''s word of the year Marketplace Tech
  5. 存储三巨头破万亿市值,存储超级周期何时能见顶?| S10E13 What's Next|科技早知道
  6. EP270 一枚芯片的漫长征途:我们离“算力自由”还有多远? Talk三联
  7. E228|谷歌TPU能撼动英伟达吗?前TPU工程师首次揭秘 硅谷101
  8. Socialists Sweep NYC, China Catches Up in Coding, AI Memory Crunch, Micron's Blowout Quarter All-In with Chamath, Jason, Sacks & Friedberg