Key Takeaways
- Nvidia GTC 2026 highlighted a shift to AI inference and revealed a massive $1 trillion hardware order backlog.
- Fast integration of Groq tech and the new Vera CPU will drastically boost processing to 1,000 tokens per second.
- Nvidia resumes H200 chip manufacturing for China and plans to return 50% of free cash flow to shareholders.
The Inference Inflection
The most prominent theme of the conference was the explosion of inference demand. Inference is the actual process of an AI model generating answers and taking action. With the rise of autonomous AI agents—programs that operate continuously and perform complex coding, searching, or organizational tasks—the need for token generation is growing exponentially. The industry has officially reached an "inference inflection" point.
The Groq Integration: A Massive Leap in Speed
In a remarkably fast turnaround, Nvidia has fully integrated the technology from its recent Groq asset acquisition into its hardware roadmap. The new Groq 3 LPX server rack is scheduled to ship in the second half of 2026.
By combining the Groq LPU with Nvidia's Vera Rubin servers, the system offloads low-latency token generation tasks to Groq while utilizing Vera Rubin for attention math and massive memory bandwidth. This powerful combination allows the system to generate up to 1,000 tokens per second, drastically improving the economics of deploying AI at a gigawatt data center scale. CEO Jensen Huang estimates that Groq systems are highly specialized and could eventually make up about 25% of a total AI data center footprint.
The Vera CPU
To support this new era of agentic AI, Nvidia showcased the Vera CPU. Built on a completely redesigned ARM architecture, this processor is specifically tailored to provide industry-leading single-thread performance, memory bandwidth, and energy efficiency under heavy loads. It is designed to be the foundational CPU for systems running continuous, heavy AI agent workloads.
$1 Trillion in Order Visibility
To underscore the sheer scale of current hardware demand, Nvidia announced that it has over $1 trillion in order visibility for its Blackwell and Rubin platforms spanning from 2025 through 2027. This figure only accounts for these specific core architectures and does not include other products like standalone CPUs, networking gear, or Groq systems, indicating that the total data center demand is likely even higher.
Global Markets and Shareholder Returns
- China Market Resurgence: Nvidia confirmed it has received licenses and purchase orders to supply H200 chips to numerous customers in China. The company is currently in the process of restarting manufacturing to fulfill this specific demand.
- Shareholder Value: Nvidia plans to return 50% of its free cash flow to shareholders in the coming year through a combination of increased stock buybacks and dividends.
