Nvidia announced that its Groq 3 LPX rack is now in full production and will be deployed later this year, following the company's $20 billion acquisition of assets from chip startup Groq in December, which marks Nvidia's largest purchase to date [1]. The Groq racks will be used alongside Nvidia's Vera central processors and Rubin graphics processors at neocloud Nebius, according to Nvidia senior director Dion Harris [1]. Each Groq 3 LPX rack contains 256 individual Groq 3 chips and can deliver 3,400 tokens per second, based on a benchmark from Artificial Analysis [1]. The Groq architecture features 500 megabytes of on-die SRAM to reduce memory bottlenecks, with the chips manufactured by Samsung, while Nvidia's GPUs are produced by Taiwan Semiconductor Manufacturing [1].
The introduction of Groq racks underscores Nvidia's focus on low-latency inference, which is increasingly important for AI agents, particularly in coding applications where responsiveness is critical [1]. Harris emphasized that this technology enables cloud companies to offer premium service tiers for latency-sensitive customers [1]. He clarified that low-latency chips like Groq are not intended to replace GPUs, which remain essential for both training and inference, but rather to optimize specific parts of AI workloads, such as the 'decode' phase [1].
Competition in the low-latency AI chip market is intensifying. Advanced Micro Devices (AMD) announced earlier this year that it would integrate its rack-scale systems with Cerebras chips, which are also focused on low-latency inference and power OpenAI's new Ultrafast mode, currently delivering 750 tokens per second [1]. In contrast, Nvidia's Groq 3 LPX racks offer a significantly higher throughput at 3,400 tokens per second [1].
Nvidia is also ramping up shipments of its Vera Rubin systems, which began production earlier this year [1]. At the unveiling of the Vera Rubin and Groq 3 LPX systems in March, Nvidia CEO Jensen Huang projected $1 trillion in cumulative sales between the current-generation Blackwell chips and the new Vera Rubin systems through 2027 [1]. Huang stated that a quarter of data center space intended for coding applications would be allocated to Groq chips, with the remainder dedicated to Vera Rubin systems [1]. Nvidia is scheduled to report earnings on Wednesday [1].
CONCLUSION
Nvidia's deployment of Groq 3 LPX racks marks a major step in its strategy to dominate the low-latency AI inference market, following its record $20 billion acquisition. With competitive throughput and a clear focus on premium, latency-sensitive applications, Nvidia is positioning itself strongly against rivals like AMD and Cerebras. The company's upcoming earnings report may provide further insight into the financial impact of these initiatives.
