Next-Generation Rack-Scale AI

NVIDIA Vera Rubin NVL72

Built for the Agentic AI Factory

72 Rubin GPUs and 36 Vera CPUs operate as one rack-scale AI supercomputer—engineered for advanced reasoning, massive mixture-of-experts models, test-time scaling, and the exploding token demand of agentic AI.

ADS integrates network, storage, power, liquid cooling, facility readiness, and deployment plan required to turn Vera Rubin into production infrastructure.

Request a Consultation

NVIDIA's Rubin Positioning

From GPU infrastructure to an engine for agentic intelligence.

NVIDIA is positioning Vera Rubin around the economics of intelligence: more reasoning, more tokens, less power per unit of work, and a single architecture that spans pretraining, post-training, test-time scaling, and real-time agentic inference.

Chip to grid. Designed as one AI supercomputer.

Vera CPU, Rubin GPU, NVLink 6, ConnectX-9, BlueField-4, and Spectrum-X are co-designed as a platform rather than assembled as independent parts.

Platform at a Glance

A Major Step Beyond Blackwell

GIGABYTE's Vera Rubin NVL72 rack combines the new Rubin compute generation with substantially higher memory bandwidth, faster scale-up interconnect, next-generation networking, and 100% liquid-cooled compute trays.

72

NVIDIA Rubin GPUs

36

NVIDIA Vera CPUs

20.7 TB

HBM4 GPU memory across the rack

1,580 TB/s

Aggregate HBM4 bandwidth stated by GIGABYTE

Why Rubin Matters

Optimized for Intelligence per Dollar and Watt

NVIDIA's launch messaging for Vera Rubin is less about peak chip specs and more about AI-factory economics—how much useful model work an infrastructure investment can produce.

Up to 10×

Lower Inference Token Cost

NVIDIA states that the Rubin platform can reduce inference token generation cost by up to 10× compared with its Blackwell platform for targeted workloads.

4× Fewer

GPUs for MoE Training

NVIDIA says Rubin can train mixture-of-experts models with one-quarter the GPUs compared with Blackwell, changing the economics of large-scale model development.

3.7×

Higher Preview MLPerf Throughput

In NVIDIA's September 2026 MLPerf Inference v6.1 preview, Vera Rubin NVL72 delivered up to 3.7× higher throughput than GB300 NVL72 on Qwen3-VL.

GIGABYTE DLB5-CB0

Third-Generation Rack-Scale Architecture

The GIGABYTE DLB5-CB0 is a 48U, direct-liquid-cooled Vera Rubin NVL72 system with 18 XN16 compute trays, nine NVIDIA NVLink switch trays, four 3RU 110 kW power shelves, dual out-of-band management switches, and an in-row CDU architecture.

Each 1U XN16-CB0-L01 compute tray contains two NVIDIA Vera CPUs and four Rubin GPUs, plus ConnectX-9 SuperNICs and BlueField-4 DPUs for scale-out networking and infrastructure services.

Vera Rubin NVL72 scales up 72 GPUs as one NVLink domain, then scales out over NVIDIA Quantum-X800 InfiniBand or Spectrum-X Ethernet.

Management

2 × out-of-band management
switches

Power

2 × 3RU 110 kW power shelves

Compute

10 × 1U trays

NVLink 6 Fabric

9 × NVLink switch trays

Compute

8 × 1U trays

Power

2 × 3RU 110 kW power shelves

Cooling

100% liquid-cooled compute;
in-row CDU

Rubin Compute With Next-Generation I/O

Each tray combines high-bandwidth HBM4, the Vera CPU, NVLink 6, 800G-class networking, local NVMe, and infrastructure acceleration in a dense 1U liquid-cooled node.

Compute & Memory

CPU

2 × NVIDIA Vera CPUs, 88 Olympus cores each

GPU

4 × NVIDIA Rubin GPUs

GPU Memory

Up to 288 GB HBM4 per GPU; up to 1,152 GB per tray

Local Storage

4 × E1.S Gen6 NVMe + 1 × E1.S Gen5 boot SSD

Networking & Cooling

SuperNIC

8 × 800 Gb/s OSFP via NVIDIA ConnectX-9

DPU

2 × 400 Gb/s QSFP via NVIDIA BlueField-4

NVLink

4 × NVIDIA NVLink Switch connectors

Cooling

100% liquid-cooled with internal manifold

Agentic AI

Designed for Workloads That Don't Stop at One Model Call

Agentic systems can reason, retrieve, invoke tools, run code, call sub-agents, and loop repeatedly before a task is finished. That changes the infrastructure requirement from simply serving models to sustaining massive volumes of tokens, context, data movement, and CPU orchestration.

Reasoning & Test-Time Scaling

Accelerate workloads that spend more compute at inference time to improve the quality of complex answers and multistep reasoning.

Agentic Inference

Support agents and sub-agents that continuously reason, retrieve information, call tools, and generate large token volumes.

Massive MoE Models

NVLink 6 and Rubin's compute architecture are positioned for training and serving increasingly large mixture-of-experts models.

Confidential AI at Rack Scale

NVIDIA positions Vera Rubin NVL72 as its first rack-scale platform extending confidential computing across CPU, GPU, and NVLink domains.

AI + HPC Convergence

Rubin also targets scientific computing, simulation, data processing, and AI-for-science workloads—not just generative AI.

Continuous Software Gains

NVIDIA's latest MLPerf results emphasize platform-level performance improvements driven by software frameworks such as Dynamo, vLLM, and TensorRT-LLM.

The ADS Value Add

Don't Just Buy a Rubin Rack. Deploy the AI Factory Around It.

At 240 kW per rack, Vera Rubin changes the infrastructure problem. ADS helps turn the rack into a deployable system by engineering the surrounding network, storage, facility, integration, validation, and support model.

The value of Vera Rubin is not just more compute. It is more intelligence per unit of power, infrastructure, and capital—and ADS helps engineer the system required to realize that value.

Workload-First Architecture

We start with the workload—reasoning, training, inference, AI for science, or mixed use—and design the surrounding infrastructure around its data and scale requirements.

High-Speed Fabric Integration

ADS can architect the scale-out fabric around NVIDIA Quantum-X800 InfiniBand or Spectrum-X Ethernet and the ConnectX-9 interfaces built into the compute trays.

Storage Matched to the Workload

Choose IBM Storage Scale, BeeGFS, VDURA, custom Ceph, or another appropriate data architecture based on throughput, checkpointing, context, object, file, and capacity needs.

240 kW Facility Readiness

Power distribution, in-row CDU capacity, water quality, heat rejection, rack placement, and network uplinks must be resolved before a rack arrives—not after.

Integration & Validation

ADS coordinates and validates the infrastructure as a system so compute, fabric, storage, cooling, and software are aligned before production workloads depend on them.

One Accountable Partner

Instead of managing a collection of vendors independently, customers get an integration partner focused on the end-to-end outcome from architecture through lifecycle support.

Data Infrastructure

Keep 72 Rubin GPUs Fed With the Right Data Architecture

The faster the compute, the more important the path to data becomes. ADS can integrate high-performance storage with the scale-out network so model data, checkpoints, datasets, vector data, context, and outputs move without becoming the next bottleneck.

Planning for Vera Rubin?

Bring ADS your model roadmap, scale target, data profile, networking preferences, power envelope, and cooling constraints. We'll help turn them into a production-ready Vera Rubin architecture.

Request a Consultation

Applied Data Systems

©2026 Applied Data Systems