Leading AI labs are currently constrained by electricity and compute costs, often paying high margins to Nvidia for model training hardware. Google bucks this trend with its upcoming eighth-generation Tensor Processing Units (TPU 8t and TPU 8i), each designed for distinct AI tasks: TPU 8t optimizes training for frontier models, while TPU 8i focuses on inference for agentic applications and real-time processing. Google’s SVP Amin Vahdat emphasized their fully integrated AI stack, allowing them to cut costs per token more effectively than competitors reliant on Nvidia chips. TPU 8t offers massive scalability, supporting over a million chips in a single job with enhancements like TPU Direct Storage to speed data access. TPU 8i, developed in partnership with DeepMind, features a redesigned network topology for drastically lower latency, improving performance for real-time language model sampling and reinforcement learning. This vertical integration helps Google avoid the high “Nvidia tax,” representing a significant cost advantage for enterprise buyers. These advances suggest a shift in cloud AI evaluations, emphasizing tailored hardware capabilities alongside raw performance metrics.
Back