Particle.news
Download on the App Store

NVIDIA Puts Groq 3 LPX Into Full Production to Speed Agentic AI

The company says the low‑latency LPX racks aim to cut token cost and power per unit of work and to make multi‑step AI agents more responsive.

Overview

  • NVIDIA announced Monday that Groq 3 LPX is in full production and will be deployed as an inference extension of its Vera Rubin NVL72 platform.
  • The company reported a 3,400 output‑tokens‑per‑second result running Gemma 4 31B with a 100,000‑token context in Artificial Analysis benchmarking, a claim that has not yet seen independent validation.
  • Nebius is named as the first cloud customer to put Groq 3 LPX into production and SpaceXAI confirmed plans to use Vera CPUs and Rubin racks for Grok with an optimized Rubin rack planned for a Starmind orbital satellite.
  • NVIDIA says Groq 3 chips are made by Samsung and that it packs 256 Groq 3 chips per LPX rack, positioning LPX as a low‑latency decode accelerator that complements GPUs rather than replacing them.
  • NVIDIA also released Vera Rubin efficiency figures claiming up to 30x higher throughput per megawatt and up to 35x lower token cost versus prior NVL72 systems and said those results are pending external review, which could affect cloud pricing and how operators build low‑latency tiers.