Skip to content

NVIDIA-Certified Hypervisors: Bringing Near Bare-Metal Performance to GPU VMs

Part 1 of a three-part series on NVIDIA-Certified Hypervisors

In August 2026, Rafay announced that its Virtual Machines-as-a-Service (VMaaS) offering had achieved NVIDIA-Certified Hypervisors status for NVIDIA accelerated computing infrastructure.

This milestone comes as enterprises, AI factories, and GPU cloud providers increasingly turn to virtualization to securely deliver GPU infrastructure across customers, teams, and workloads. Virtual machines provide the isolation, multi-tenancy, governance, and operational flexibility needed to turn high-value GPU infrastructure into a scalable cloud service.

But for AI workloads, virtualization raises a critical question:

Can a GPU VM deliver performance comparable to running directly on bare metal?

The NVIDIA-Certified Hypervisors program is designed to answer that question by validating virtualization platforms for performance on NVIDIA accelerated computing infrastructure.

In this three-part series, we'll look beyond the certification itself and explore what NVIDIA validates, how the testing is performed, and what the results tell us about the performance of GPU workloads running inside virtual machines.

Hypervisor Certification


What Are NVIDIA-Certified Hypervisors?

The NVIDIA-Certified Hypervisors program validates virtualization platforms designed for GPU-accelerated AI and compute workloads. NVIDIA describes certified hypervisors as performance-optimized virtualization platforms for enterprise AI data centers and AI factories.

Certification testing evaluates representative performance-critical behaviors across areas such as GPU compute, memory, data-path efficiency, inter-GPU communication, high-speed GPU networking, and LLM inference. By accurately exposing the underlying hardware topology and implementing the necessary performance optimizations, certified platforms are designed to deliver near bare-metal performance for representative AI and accelerated computing workloads.

This makes certification particularly relevant for AI factories, NeoClouds, NVIDIA Cloud Partners, and enterprises that want the operational advantages of virtualization without significantly compromising the performance of their GPU infrastructure.


Arm Platforms

The Arm platform certification is currently scoped to NVIDIA GB200 NVL systems, which combine NVIDIA Blackwell GPUs with NVIDIA Grace CPUs as part of NVIDIA's rack-scale architecture. This track validates the performance of a single 4-GPU, 2-CPU passthrough VM within one GB200 NVL compute tray.

The certification tests an important deployment model for Blackwell infrastructure: exposing the CPU and GPU resources of a GB200 compute tray directly to a VM while retaining the isolation and lifecycle benefits of virtualization.


x86 Platforms

The x86 platform certification covers systems built on the NVIDIA HGX platform as well as NVIDIA DGX systems. Certification is currently scoped to NVSwitch-based, 8-GPU NVIDIA Hopper SXM systems and consists of two required phases.


Summary

The certification tracks can be summarized as follows:

Track Platform VM Configuration Scope
Arm NVIDIA GB200 NVL 4 GPUs + 2 CPUs Single VM within one GB200 NVL compute tray
x86 – Single-Node NVIDIA Hopper SXM 8 GPUs Single VM on one node; full GPU/NIC passthrough
x86 – Multi-Node NVIDIA Hopper SXM 2 × 8-GPU VMs Two VMs across two identical nodes

Together, these tracks validate that GPU virtualization can provide the isolation and operational flexibility of virtual machines while delivering performance approaching bare metal.


Upcoming Blogs

What does it actually take to demonstrate that a GPU VM can deliver near bare-metal performance?That is exactly what we'll examine in the next two posts in this series.

Part 2 will be focused on test results performed on a single 8-GPU VM and examine the results across GPU compute, memory, GPU-to-GPU communication, networking, and AI inference.

Part 3 will be focused on a test environment spanning two GPU systems and examine how virtualization performs for distributed workloads where GPU-to-GPU communication and high-speed networking between nodes become critical.

Together, the results provide a practical look at what near bare-metal GPU performance inside virtual machines actually means—and why it matters for organizations building production AI clouds and AI factories.