Next-Generation Rack-Scale AI
NVIDIA Vera Rubin NVL72
Built for the Agentic AI Factory
72 Rubin GPUs and 36 Vera CPUs operate as one rack-scale AI supercomputer—engineered for advanced reasoning, massive mixture-of-experts models, test-time scaling, and the exploding token demand of agentic AI.
ADS integrates network, storage, power, liquid cooling, facility readiness, and deployment plan required to turn Vera Rubin into production infrastructure.
Request a Consultation

NVIDIA's Rubin Positioning
From GPU infrastructure to an engine for agentic intelligence.
NVIDIA is positioning Vera Rubin around the economics of intelligence: more reasoning, more tokens, less power per unit of work, and a single architecture that spans pretraining, post-training, test-time scaling, and real-time agentic inference.
Chip to grid. Designed as one AI supercomputer.
Vera CPU, Rubin GPU, NVLink 6, ConnectX-9, BlueField-4, and Spectrum-X are co-designed as a platform rather than assembled as independent parts.
Platform at a Glance
A Major Step Beyond Blackwell
GIGABYTE's Vera Rubin NVL72 rack combines the new Rubin compute generation with substantially higher memory bandwidth, faster scale-up interconnect, next-generation networking, and 100% liquid-cooled compute trays.
72
NVIDIA Rubin GPUs
36
NVIDIA Vera CPUs
20.7 TB
HBM4 GPU memory across the rack
1,580 TB/s
Aggregate HBM4 bandwidth stated by GIGABYTE
Why Rubin Matters
Optimized for Intelligence per Dollar and Watt
NVIDIA's launch messaging for Vera Rubin is less about peak chip specs and more about AI-factory economics—how much useful model work an infrastructure investment can produce.
Up to 10×
Lower Inference Token Cost
NVIDIA states that the Rubin platform can reduce inference token generation cost by up to 10× compared with its Blackwell platform for targeted workloads.
4× Fewer
GPUs for MoE Training
NVIDIA says Rubin can train mixture-of-experts models with one-quarter the GPUs compared with Blackwell, changing the economics of large-scale model development.
3.7×
Higher Preview MLPerf Throughput
In NVIDIA's September 2026 MLPerf Inference v6.1 preview, Vera Rubin NVL72 delivered up to 3.7× higher throughput than GB300 NVL72 on Qwen3-VL.
GIGABYTE DLB5-CB0
Third-Generation Rack-Scale Architecture
The GIGABYTE DLB5-CB0 is a 48U, direct-liquid-cooled Vera Rubin NVL72 system with 18 XN16 compute trays, nine NVIDIA NVLink switch trays, four 3RU 110 kW power shelves, dual out-of-band management switches, and an in-row CDU architecture.
Each 1U XN16-CB0-L01 compute tray contains two NVIDIA Vera CPUs and four Rubin GPUs, plus ConnectX-9 SuperNICs and BlueField-4 DPUs for scale-out networking and infrastructure services.
Vera Rubin NVL72 scales up 72 GPUs as one NVLink domain, then scales out over NVIDIA Quantum-X800 InfiniBand or Spectrum-X Ethernet.
Management
2 × out-of-band management
switches
Power
2 × 3RU 110 kW power shelves
Compute
10 × 1U trays
NVLink 6 Fabric
9 × NVLink switch trays
Compute
8 × 1U trays
Power
2 × 3RU 110 kW power shelves
Cooling
100% liquid-cooled compute;
in-row CDU

Rubin Compute With Next-Generation I/O
Each tray combines high-bandwidth HBM4, the Vera CPU, NVLink 6, 800G-class networking, local NVMe, and infrastructure acceleration in a dense 1U liquid-cooled node.
Compute & Memory
CPU
2 × NVIDIA Vera CPUs, 88 Olympus cores each
GPU
4 × NVIDIA Rubin GPUs
GPU Memory
Up to 288 GB HBM4 per GPU; up to 1,152 GB per tray
Local Storage
4 × E1.S Gen6 NVMe + 1 × E1.S Gen5 boot SSD
Networking & Cooling
SuperNIC
8 × 800 Gb/s OSFP via NVIDIA ConnectX-9
DPU
2 × 400 Gb/s QSFP via NVIDIA BlueField-4
NVLink
4 × NVIDIA NVLink Switch connectors
Cooling
100% liquid-cooled with internal manifold
Agentic AI
Designed for Workloads That Don't Stop at One Model Call
Agentic systems can reason, retrieve, invoke tools, run code, call sub-agents, and loop repeatedly before a task is finished. That changes the infrastructure requirement from simply serving models to sustaining massive volumes of tokens, context, data movement, and CPU orchestration.
Reasoning & Test-Time Scaling
Accelerate workloads that spend more compute at inference time to improve the quality of complex answers and multistep reasoning.

Agentic Inference
Support agents and sub-agents that continuously reason, retrieve information, call tools, and generate large token volumes.

Massive MoE Models
NVLink 6 and Rubin's compute architecture are positioned for training and serving increasingly large mixture-of-experts models.

Confidential AI at Rack Scale
NVIDIA positions Vera Rubin NVL72 as its first rack-scale platform extending confidential computing across CPU, GPU, and NVLink domains.

AI + HPC Convergence
Rubin also targets scientific computing, simulation, data processing, and AI-for-science workloads—not just generative AI.

Continuous Software Gains
NVIDIA's latest MLPerf results emphasize platform-level performance improvements driven by software frameworks such as Dynamo, vLLM, and TensorRT-LLM.

The ADS Value Add
Don't Just Buy a Rubin Rack. Deploy the AI Factory Around It.
At 240 kW per rack, Vera Rubin changes the infrastructure problem. ADS helps turn the rack into a deployable system by engineering the surrounding network, storage, facility, integration, validation, and support model.
The value of Vera Rubin is not just more compute. It is more intelligence per unit of power, infrastructure, and capital—and ADS helps engineer the system required to realize that value.
Workload-First Architecture
We start with the workload—reasoning, training, inference, AI for science, or mixed use—and design the surrounding infrastructure around its data and scale requirements.
High-Speed Fabric Integration
ADS can architect the scale-out fabric around NVIDIA Quantum-X800 InfiniBand or Spectrum-X Ethernet and the ConnectX-9 interfaces built into the compute trays.
Storage Matched to the Workload
Choose IBM Storage Scale, BeeGFS, VDURA, custom Ceph, or another appropriate data architecture based on throughput, checkpointing, context, object, file, and capacity needs.
240 kW Facility Readiness
Power distribution, in-row CDU capacity, water quality, heat rejection, rack placement, and network uplinks must be resolved before a rack arrives—not after.
Integration & Validation
ADS coordinates and validates the infrastructure as a system so compute, fabric, storage, cooling, and software are aligned before production workloads depend on them.
One Accountable Partner
Instead of managing a collection of vendors independently, customers get an integration partner focused on the end-to-end outcome from architecture through lifecycle support.
Data Infrastructure
Keep 72 Rubin GPUs Fed With the Right Data Architecture
The faster the compute, the more important the path to data becomes. ADS can integrate high-performance storage with the scale-out network so model data, checkpoints, datasets, vector data, context, and outputs move without becoming the next bottleneck.

Planning for Vera Rubin?
Bring ADS your model roadmap, scale target, data profile, networking preferences, power envelope, and cooling constraints. We'll help turn them into a production-ready Vera Rubin architecture.
Request a Consultation

Applied Data Systems
©2026 Applied Data Systems
