Nvidia’s Groq 3 LPX Inference Chip: A New Era of AI Acceleration

Nvidia has reached a significant milestone with the full production of its dedicated AI inference accelerator, the Groq 3 LPX. This announcement, made at Hot Chips 2026 on August 24, 2026, marks a major acceleration in the availability of high-speed AI inference hardware for agentic workloads.

The Groq 3 LPX is a powerful extension of Nvidia’s flagship Vera Rubin data center platform. It’s specifically designed to accelerate the “decode” phase of AI inference, which determines the speed at which individual users receive token-by-token responses from large language models.

Nvidia incorporates 256 individual Groq 3 chips into its LPX racks. The chip, fabricated by Samsung, integrates 500 megabytes of SRAM directly onto the die to eliminate memory bandwidth bottlenecks. In benchmarking by Artificial Analysis, the Groq 3 LPX showcased an impressive 3,400 output tokens per second running the Gemma 4 31B model with a 100,000-token context window.

Neocloud provider Nebius Group has already become the chip’s first committed commercial customer. Nvidia strategically positions the Groq 3 LPX as a complement to its GPU lineup, not a replacement. The chips collaborate to handle different parts of the AI workload pipeline. This production launch is conveniently timed, coming just before Nvidia’s quarterly earnings release on Wednesday.

Source: SiliconANGLE – Nvidia’s dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents

Move to the category:

Leave a Reply

Your email address will not be published. Required fields are marked *