NVIDIA says Groq 3 LPX is in full production for agentic inference
At Hot Chips, NVIDIA called Groq 3 LPX an interactive inference accelerator that extends Vera Rubin. On Gemma 4 31B with 100k context, Artificial Analysis measured 3,400 output tokens per second.

On August 24 NVIDIA said at Hot Chips that the Groq 3 LPX interactive inference accelerator, which extends Vera Rubin, is in full production. Artificial Analysis measured Gemma 4 31B at 100k context at 3,400 output tokens per second — four times the nearest alternative.
What we know
- Groq 3 LPX is described as an interactive inference accelerator extending Vera Rubin and is now in full production.
- Artificial Analysis recorded 3,400 output tokens per second on Gemma 4 31B with 100k context, four times the nearest alternative.
- Nebius is the first AI cloud; Groq’s own cloud is among the earliest adopters.
- CNBC reported separately that NVIDIA bought Groq assets for $20 billion in December and that each LPX rack holds 256 Groq 3 chips.
Takeaways
- NVIDIA is selling Groq silicon as a production inference add-on to Vera Rubin, not a lab demo.
- The headline benchmark is decode speed on a 31B model at long context.
- The $20 billion December asset purchase and 256-chip racks come from CNBC, not the NVIDIA note.
Source: NVIDIA


