SecurityBrief US - Technology news for CISOs & cybersecurity decision-makers
United States
NVIDIA puts Vera Rubin into production for AI inference

NVIDIA puts Vera Rubin into production for AI inference

Tue, 25th Aug 2026 (Today)
Sean Mitchell
SEAN MITCHELL Publisher

NVIDIA has put its Vera Rubin NVL72 system with Groq 3 LPX into full production, targeting AI inference for agentic systems.

The announcement adds an inference component to NVIDIA's broader Vera Rubin platform, which is positioned around long-context processing, token generation and networking for large AI deployments.

In a benchmark by Artificial Analysis using the Gemma 4 31B model, the system produced 3,400 output tokens per second for 100,000-token context workloads, according to NVIDIA. NVIDIA said that was four times faster than the nearest alternative platform in the test.

NVIDIA is responding to a shift in AI infrastructure demand as customers move from model training to inference and reasoning workloads. Those tasks increasingly involve systems that process larger context windows, generate responses token by token and interact with other software tools or AI models.

To support that shift, NVIDIA is pairing its Rubin GPUs with Groq 3 LPX for token generation. The design separates large-scale context processing from latency-sensitive decoding, which NVIDIA says is intended to improve responsiveness in agent-based applications.

Partner uptake

Several customers and partners are adopting parts of the Vera Rubin platform.

Nebius is the first AI cloud provider to adopt NVIDIA Groq 3 LPX, according to NVIDIA. The addition to its Nebius Token Factory is intended to increase inference speeds for developers building interactive AI agents, coding systems and other real-time services.

CoreWeave has deployed Spectrum-X Multiplane in production to connect Vera Rubin racks across its cloud infrastructure. NVIDIA said the setup uses multiple parallel switches to create a flatter network while avoiding the added latency and cost of a third network tier.

SpaceXAI is also adopting NVIDIA Vera CPUs as part of its next AI architecture. NVIDIA said the processors are intended for CPU-heavy agentic AI work, including orchestration, tool use, code execution, data processing and simulation, and that the systems are planned for use in data centres and orbital satellites.

Network changes

NVIDIA is also expanding the networking layer around Vera Rubin. Spectrum-X Multiplane splits each server connection into several independent paths, or planes, with each plane running a separate two-tier network.

The design allows networks to scale to 512,000 GPUs without adding a third tier, according to NVIDIA. It also said that in an eight-plane topology, the network can retain about 90% of total bandwidth if one plane fails, with hardware recovery 11 times faster than software-based multiplane load balancing.

NVIDIA said Spectrum-X Ethernet is designed to deliver 1.6 times better AI networking performance than standard Ethernet in AI infrastructure. It added that the same architecture can be extended across data centres through Spectrum-XGS Ethernet, which it said improves multi-site NCCL collectives by 1.9 times.

Scale-In launch

Alongside the production launch, NVIDIA introduced Scale-In, a new networking and infrastructure layer for AI factories based on BlueField-4 processors and the DOCA software platform.

The product is intended to handle multi-tenant networking, storage access, security, provisioning and observability without relying on host compute resources. NVIDIA is presenting it as a way to move infrastructure services closer to the AI system itself as shared AI environments grow larger and more complex.

NVIDIA is also broadening its approach to custom silicon through NVLink Fusion, which links custom XPUs and CPUs into its scale-up and scale-out infrastructure. That gives hyperscalers and AI-focused companies a way to combine in-house chips with NVIDIA networking, rack designs and management software.

By standardising systems around a shared architecture, operators can use the same rack footprint, networking, cooling and power systems while changing the mix of GPUs and XPUs over time, NVIDIA said. The approach reflects the company's effort to maintain its position not only in AI chips but across the wider infrastructure stack that supports model deployment.

The latest product moves show NVIDIA pushing tighter integration of compute, inference and networking as customers look to handle larger numbers of AI queries with lower delay. Rather than selling a single processor as the centre of the system, the company is trying to tie together the rack, interconnect and inference layer as one platform for running agent-based AI at scale.

A rack-scale NVIDIA Groq 3 LPX deployment can include 256 LP30 accelerators connected through direct chip-to-chip links, creating what NVIDIA describes as a deterministic inference engine for modern AI factories.