Computing

Google is building a server chip it claims is Up to 10 times more efficient than its TPUs

(2 months ago) · 2 min read · By Future Technology · Edited by Nath Connell

Every big AI company has arrived at the same conclusion from a different direction: the model is not the expensive part any more, serving it is. Google appears to be answering that with a new server chip, reportedly code-named Frozen v2, and the number attached to it is the reason people are paying attention. Internal sources put its efficiency at 6 to 10 times that of Google's current TPUs.

The interesting claim is not the raw speed, it is the design approach. Frozen v2 is described as being built around the Gemini architecture rather than being a general accelerator that Gemini happens to run on. If you know exactly which model shapes you need to serve, you can throw away a lot of flexibility and spend that silicon on the operations you actually use. That is the same logic behind Amazon's Inferentia and OpenAI's reported in-house chip work.

The caution is straightforward. Google has not officially confirmed the chip or the performance figures, and efficiency numbers quoted before a part ships in high volume almost always come from a favourable benchmark. Yields, memory bandwidth and the messy reality of a datacentre tend to trim the headline. Treat 6 to 10 times as the ceiling of a marketing case, not the floor of a spec sheet.

Still, the direction of travel is clear enough. The companies with the biggest inference bills are all trying to stop renting margin to Nvidia, and each one that succeeds pulls a little demand out of the merchant chip market. Whether Frozen v2 hits its numbers matters less than the fact that Google thinks building it is worth the effort.

[### Tesla and SpaceX Break Ground on Terafab, a 16.8 Billion Dollar Texas Chip Plant

Tesla and SpaceX are building Terafab, a 16.8bn dollar Texas chip factory to feed their own AI, robotics and space compu](/article/tesla-spacex-terafab-texas-chip-plant/)[### NVIDIA Vera Rubin NVL72 Ramps Up at CoreWeave and Google

NVIDIA's Vera Rubin NVL72, the company's next-generation GPU architecture, is now in active production and running at Co](/nvidia-vera-rubin-nvl72-production-ramp/)

Browse all Hardware stories →

More from Future Technology