
Pinterest is using the NVIDIA Blackwell platform and NVIDIA Dynamo inference software to bring conversational AI to visual discovery.
The news comes as agentic AI is driving a new class of workloads that demand more performance, efficiency and scale from AI infrastructure.
NVIDIA addresses that challenge with a full-stack AI factory platform spanning Vera Rubin systems, Dynamo inference software, NeMo libraries and NVIDIA networking — including NVIDIA NVLink for scale-up computing, Spectrum-X Ethernet and ConnectX SuperNICs for connecting thousands of nodes, BlueField-powered context-memory storage and BlueField DPUs for infrastructure security.
“Infrastructure that’s fungible, that’s reliable, that’s going to last 10 years and really becomes an asset for computing the world’s computing problems in the world’s industries — they can build that with Vera Rubin, with DSX and all the innovations that we have here,” said Buck.
The metric for AI infrastructure is fast shifting from peak performance to validated agentic tokens per megawatt. AI factories must now be codesigned from silicon to grid. NVIDIA DSX MaxLPS can deliver up to 1.4x more tokens per megawatt through factory-wide power optimization, while NVLink helps unite large-scale accelerated computing into a single high-performance system.
The result is AI infrastructure designed to generate more tokens, improve efficiency and help customers get more value from every megawatt of power.
AI inference economics increasingly come down to three factors: system performance, infrastructure scaling and software optimization. Together, they determine how many tokens can be generated, how efficiently AI services scale and how much value organizations extract from their infrastructure investments.
The NVIDIA platform is designed to optimize across all three, while giving enterprises the flexibility to run any model and workload on the same infrastructure, from training and inference to recommender systems, reasoning models and generative AI.
New results from MLPerf Inference v6.1 — the MLCommons consortium’s long-standing industry benchmark, spanning a broad range of models with every result peer-reviewed before publication — highlight that advantage. In its first MLPerf Inference preview submission, the NVIDIA Vera Rubin NVL72 system delivered up to 3.7x higher throughput than NVIDIA GB300 NVL72, demonstrating the performance gains possible with NVIDIA’s next-generation AI infrastructure.
NVIDIA GB300 NVL72 also showcased industry-leading scalability. A 288-GPU submission spanning four GB300 NVL72 racks achieved 99% scaling efficiency, with throughput growing nearly linearly from a single-rack baseline.
The results also underscore the impact of software innovation. NVIDIA’s MLPerf Inference v6.1 submissions delivered up to 1.6x higher performance than v6.0 through software optimizations alone, with additional gains achieved after the benchmark submission period.
For organizations evaluating AI infrastructure, performance, scaling efficiency and software velocity remain critical drivers of long-term inference economics.
Learn more about the NVIDIA Vera Rubin NVL72 benchmark results.
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://blogs.nvidia.com/blog/ai-infra-summit-vera-rubin-dsx-energy-efficiencies-tokens-per-watt-ai-factories/#primary
- https://blogs.nvidia.com/blog/author/nvidiawriters/
- https://blogs.nvidia.com/blog/ai-infra-summit-vera-rubin-dsx-energy-efficiencies-tokens-per-watt-ai-factories/#disqus_thread
- SK hynix reportedly discussing US memory chip manufacturing with Intel — options include leasing Ohio plant or forming joint venture with other AI hyperscalers
- Perplexity’s local AI agent comes to Windows, but only for RTX GPUs with at least 24GB of VRAM — Portable Computer brings AI for multistep tasks to compatible P
- Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX
- Valve engineers discuss the duality of the Steam Frame and pricing — Valve's newest VR headset pivots SteamOS to Arm
- ChatGPT transcripts are reportedly read by humans to improve responses, including those with personal information — 'Project Lilly' has seen OpenAI hire hundred
Informational only. No financial advice. Do your own research.