Overview
- Multiple reports say Google is building a server chip codenamed Frozen v2 that hard‑codes Gemini’s underlying compute architecture into silicon to speed and slim down inference.
- Insiders quoted by outlets estimate the design could boost tokens‑per‑watt roughly six to ten times versus Google’s current in‑house chips, though those efficiency figures are reported and not independently confirmed.
- The project is described as a response to internal AI compute shortages that have strained Google Cloud capacity and reportedly forced the company to turn down some customer orders.
- Frozen v2 evolved from an earlier Jeff Dean‑led idea to burn full model weights into silicon; the new approach fixes the architecture in hardware but keeps model weights updateable to avoid locking chips to a single Gemini version.
- Google frames Frozen v2 as an exploratory, limited‑scale product line that would sit alongside TPUs rather than replace them, and the company has not committed to large‑scale mass production.