entity Updated 2026-08-07 Tags: Ai, Chip, Google, Tpu, Inference

Ironwood TPU

Ironwood TPU is the TPU generation discussed in E228|谷歌TPU能撼动英伟达吗?前TPU工程师首次揭秘. [[HenryTPUEngineer|Henry]] describes Ironwood V7 as having meaningfully improved peak FLOPS and memory bandwidth, and says it is aimed especially at [[HighThroughputInferenceBatching|large-scale inference]] needs such as lower latency, higher throughput, and enough [[HighBandwidthMemory|memory bandwidth]] for LLM decode.

In the source, Ironwood matters less as a standalone chip benchmark than as evidence of Google shifting TPU design toward the inference-heavy economics of deployed Gemini and cloud workloads. Its value still depends on [[XLACompiler|XLA]], [[TPUPodSystemOptimization|TPU Pod]] design, Advanced Packaging, supply ramp, and integration into Google Cloud.

Connections