Token per Watt
Token per watt is the efficiency metric foregrounded in E230|1万亿收入预期背后:英伟达的巅峰与软肋 through Nvidia’s GTC messaging. The episode treats it as a sign that AI infrastructure is being measured by usable token output per energy input, not only by FLOPS, chip count, or benchmark speed.
The metric matters because AI Inference Cost Structure increasingly depends on power, cooling, memory movement, and data-center deployment. In the source, [[NvidiaBlackwellPlatform|Blackwell]] and [[NvidiaVeraRubinPlatform|Vera Rubin]] efficiency claims only become useful if they translate into lower cost and higher throughput inside real MaaS Infrastructure environments.
Key Claims
- Token-per-watt shifts attention from raw accelerator specs to delivered AI work under energy constraints.
- The metric links model-serving economics to [[DataCenterPowerBottleneck|data-center power bottlenecks]] and Data Center Thermal Management.
- Communication and memory movement can reduce effective token-per-watt even when arithmetic units are fast.
- Token efficiency can increase total demand if it enables more agents and applications to run continuously.
Connections
- Nvidia, GPU, Nvidia Blackwell Platform, and Nvidia Vera Rubin Platform - platform and product context.
- AI Inference Cost Structure, Inference as Cash Flow, and Jevons Paradox In AI - demand and cost interpretation.
- Data Center Power Bottleneck, Data Center Thermal Management, and MaaS Infrastructure - physical serving constraints.