Particle.news
Download on the App Store

Google Developing Gemini‑Tuned Frozen v2 Chip to Sharply Cut Inference Power

The experimental server chip would embed Gemini's low‑level compute design in silicon with support for updating model weights to boost per‑watt token throughput.

Overview

  • Multiple reports say Google is building a server chip codenamed Frozen v2 that hard‑codes Gemini’s underlying compute architecture into silicon to speed and slim down inference.
  • Insiders quoted by outlets estimate the design could boost tokens‑per‑watt roughly six to ten times versus Google’s current in‑house chips, though those efficiency figures are reported and not independently confirmed.
  • The project is described as a response to internal AI compute shortages that have strained Google Cloud capacity and reportedly forced the company to turn down some customer orders.
  • Frozen v2 evolved from an earlier Jeff Dean‑led idea to burn full model weights into silicon; the new approach fixes the architecture in hardware but keeps model weights updateable to avoid locking chips to a single Gemini version.
  • Google frames Frozen v2 as an exploratory, limited‑scale product line that would sit alongside TPUs rather than replace them, and the company has not committed to large‑scale mass production.