Nvidia's Groq 3 LPX inference chip enters full production, commercializing its $20 billion Groq acquisition; racks to go live this year at Nebius

Nvidia has announced that its Groq 3 LPX inference chip has entered full production, marking the commercialization of its $20 billion acquisition of Groq. The chip, designed to enhance AI inference speed, will debut at Nebius later this year, featuring 256 chips per rack and impressive performance metrics.

qz.com qz.com+2 sources24 August 2026 · 18:38 UTC
CuriousCats Full Story

Nvidia's Groq 3 LPX inference chip has officially entered full production, a significant step following the company's $20 billion acquisition of Groq. This chip is designed to accelerate the decode phase of AI inference, crucial for generating tokens quickly for users.156

The Groq 3 LPX is integrated with Nvidia's Vera Rubin platform and is expected to deliver 3,400 output tokens per second, making it four times faster than its nearest competitor for latency-sensitive workloads. Nvidia senior director Dion Harris emphasized that the chip allows cloud providers to offer premium service tiers, stating, “For folks who are serving tokens, it unlocks the ability to offer premium tiers of service for those users and those customers who actually demand the most latency-sensitive service agreements.”

The first deployment of the Groq 3 LPX will be at Nebius, with racks set to go live later this year. Nebius's chief technology officer, Danila Shtan, noted, “Generation is the phase of inference that determines how responsive an AI system actually is, and that's exactly what NVIDIA Groq 3 LPX is built to accelerate.”23

The chip's architecture includes 500 megabytes of SRAM directly on the die, which helps mitigate memory bandwidth issues that can hinder performance. This design allows for effective tensor parallelism, essential for handling large models with high interactivity.

Overall, Nvidia's push to manufacture the Groq 3 LPX highlights the increasing demand for low-latency inference in AI applications, particularly in coding environments.

Key Insight
“The chip packs 500 megabytes of SRAM on the die to avoid memory bottlenecks, and benchmarks show 3,431 output tokens per second on Gemma 4 31B with a 100K context. Nvidia positions it as a complement to GPUs, not a replacement, targeting premium latency-sensitive service tiers.”
CuriousCats studied:
1
qz.comqz.com
“Nvidia that its Groq 3 LPX inference chip has entered full production — a milestone that brings to market the technology behind Nvidia's $20 billion purchase of chip startup Groq's assets in December, the largest deal the company has ever closed.”
qz.com →
2
CNBCCNBC
“Nvidia announced Monday that its Groq 3 LPX rack is in full production, marking the commercialization of technology from the company's largest acquisition on record.”
CNBC →
3
NVIDIA DeveloperNVIDIA Developer
“NVIDIA Groq 3 LPX, integrated with the NVIDIA Vera Rubin NVL72 platform, achieved a world-class 3,431 output tokens/second on the Artificial Analysis 100K context benchmark with the Gemma 4 31B model, demonstrating leading high-interactivity performance at long context lengths without loss in precision or model quality.”
NVIDIA Developer →
Ask CuriousCats
What is Nvidia's Groq 3 LPX chip?
Why did Nvidia acquire Groq for $20 billion?
How does this chip address memory bottlenecks?
Are other companies developing competing inference chips?
How does Groq 3 LPX perform compared to GPUs?
Get your CIA-level briefing,
in real time.
CuriousCats monitors the internet every minute for you and brings you the most personalized brief of videos, social media posts, news and more.
Download the App
Liked the depth here?
Get the full internet briefed for you any time of the day.
Get CuriousCats