Ironwood TPU
Ironwood TPU is the TPU generation discussed in E228|谷歌TPU能撼动英伟达吗?前TPU工程师首次揭秘. [[HenryTPUEngineer|Henry]] describes Ironwood V7 as having meaningfully improved peak FLOPS and memory bandwidth, and says it is aimed especially at [[HighThroughputInferenceBatching|large-scale inference]] needs such as lower latency, higher throughput, and enough [[HighBandwidthMemory|memory bandwidth]] for LLM decode.
In the source, Ironwood matters less as a standalone chip benchmark than as evidence of Google shifting TPU design toward the inference-heavy economics of deployed Gemini and cloud workloads. Its value still depends on [[XLACompiler|XLA]], [[TPUPodSystemOptimization|TPU Pod]] design, Advanced Packaging, supply ramp, and integration into Google Cloud.
Connections
- TPU, Google, Gemini, and Google Cloud — platform context.
- Memory Wall, High Bandwidth Memory, and AI Inference Cost Structure — inference bottleneck and cost context.
- AI Chip Specialization and ASIC Workload Prediction Risk — specialization and model-churn tradeoff.