Particle.news
Download on the App Store

NVIDIA Moves Groq 3 LPX Into Full Production and Secures Major Vera Rubin Orders

The production ramp signals faster, cheaper token generation for agentic AI with large commercial and orbital commitments following.

Overview

  • NVIDIA said Monday that Groq 3 LPX racks are in full production and will be deployed later this year, citing an Artificial Analysis benchmark of 3,400 output tokens per second for latency‑sensitive decode work.
  • NVIDIA published vendor-run results for its Vera Rubin NVL72 showing up to 30x higher throughput per megawatt and up to 35x lower token cost on the SemiAnalysis AgentX workload, but those numbers await broader independent verification.
  • SpaceXAI committed to use NVIDIA Vera CPUs and Vera Rubin infrastructure for Grok and announced a space‑optimized Vera Rubin NVL72 Starmind satellite targeted for launch in Q4 2027, establishing a common ground-to-orbit software and hardware stack.
  • Commercial demand is already large: AM Intelligence placed a binding order for 9,000 Vera Rubin systems for southern India and Nebius said it will be an early cloud adopter using Groq 3 LPX for a token-focused “Token Factory” service.
  • Near-term capacity will be constrained because Groq chips are manufactured by Samsung, NVIDIA GPUs by TSMC, NVIDIA packages 256 Groq 3 chips per LPX rack, and hyperscaler demand has produced multiyear backlogs for high-end clusters.