- Nvidia's Groq 3 LPX inference chip has entered full production, marking the commercialization of technology from the company's $20 billion acquisition of Groq, its largest purchase on record.
- Nebius is the first cloud provider to deploy the Groq 3 LPX chip, with racks set to go live later this year alongside Vera Rubin systems.
- The Groq 3 LPX chip was benchmarked by Artificial Analysis at 3,431 output tokens per second on the Gemma 4 31B model with a 100K context window.
- Nvidia's acquisition of Groq for $20 billion is the company's largest purchase to date, aimed at enhancing its capabilities in AI inference technology.
- The Groq 3 LPX is designed to accelerate the decode phase of AI inference, which is crucial for generating tokens quickly for users.
Nvidia's Groq 3 LPX inference chip has officially entered full production, a significant step following the company's $20 billion acquisition of Groq. This chip is designed to accelerate the decode phase of AI inference, crucial for generating tokens quickly for users.156
The Groq 3 LPX is integrated with Nvidia's Vera Rubin platform and is expected to deliver 3,400 output tokens per second, making it four times faster than its nearest competitor for latency-sensitive workloads. Nvidia senior director Dion Harris emphasized that the chip allows cloud providers to offer premium service tiers, stating, “For folks who are serving tokens, it unlocks the ability to offer premium tiers of service for those users and those customers who actually demand the most latency-sensitive service agreements.”

The first deployment of the Groq 3 LPX will be at Nebius, with racks set to go live later this year. Nebius's chief technology officer, Danila Shtan, noted, “Generation is the phase of inference that determines how responsive an AI system actually is, and that's exactly what NVIDIA Groq 3 LPX is built to accelerate.”23

The chip's architecture includes 500 megabytes of SRAM directly on the die, which helps mitigate memory bandwidth issues that can hinder performance. This design allows for effective tensor parallelism, essential for handling large models with high interactivity.
Overall, Nvidia's push to manufacture the Groq 3 LPX highlights the increasing demand for low-latency inference in AI applications, particularly in coding environments.
“The chip packs 500 megabytes of SRAM on the die to avoid memory bottlenecks, and benchmarks show 3,431 output tokens per second on Gemma 4 31B with a 100K context. Nvidia positions it as a complement to GPUs, not a replacement, targeting premium latency-sensitive service tiers.”
