Skip to content

Why Bare-Metal Telemetry and Rack-Level Visibility Are Non-Negotiable

The rise of AI-centric compute providers (aka neoclouds) has fundamentally shifted data center management requirements. Unlike traditional cloud hyperscalers built around virtualized, lower-density web applications, neoclouds orchestrate ultra-high-density GPU clusters pushed to their physical thermal and electrical limits.

At this scale, managing bare-metal infrastructure requires moving beyond basic hypervisor dashboards to deep, real-time physical rack-level telemetry. Software solutions that unify physical data center infrastructure management (DCIM) into a single operational plane are becoming essential architecture components for AI infrastructure providers.

Rack Visibility Rafay


The Physical Realities of Modern AI Racks

NVIDIA's current rack-scale platforms illustrate just how far the numbers have moved. A single NVIDIA GB200 NVL72 rack packs 72 Blackwell GPUs and 36 Grace CPUs into one liquid-cooled cabinet drawing roughly 120 kW. This is an order of magnitude beyond a typical air-cooled enterprise rack. Its successor, the NVIDIA GB300 NVL72, pushes things further. It has 72 Blackwell Ultra GPUs paired with 36 Grace CPUs, 20 TB of pooled HBM3E memory with a nominal draw of 132–142 kW (peaking near 155 kW) in the same footprint.

These aren't edge cases. These have become the baseline unit of deployment for AI infrastructure providers.

This massive power draw brings specific operational challenges that traditional software stacks were never designed to solve:

1. Direct-to-chip Liquid Cooling

Air cooling is no longer sufficient for racks like these. The GB200 NVL72 relies on an in-rack coolant distribution unit (CDU) delivering direct-to-chip liquid cooling across its GPUs, CPUs, and NVLink switches. The GB300 NVL72 goes a step further, rejecting roughly 90% of heat to liquid and only 10% to air, within an ASHRAE W45 supply-water range.

Operating fleets of these racks means continuously tracking coolant flow rates, supply/return temperatures, and instant leak-detection sensors to protect multi-million-dollar server cabinets.

2. Granular Thermal Telemetry

Individual GPU and CPU core temperatures must be monitored continuously alongside ambient environmental conditions to prevent thermal throttling before training jobs stall or fail.

Across 72 GPUs per rack, that's telemetry at a density most legacy DCIM tools were never built to display, let alone alert on.

3. Remote Power Management

When operating bare-metal server instances at 120–150+ kW per rack, engineers need granular power control (full rack, domain, compute node) to execute hard resets, cycle isolated hardware, or handle emergency power-offs without needing on-site technicians.


Key Pillars of a Modern Neocloud Control Plane

To operate efficiently, a neocloud infrastructure management system such as the Rafay GPU Orchestration Platform must bridge the gap between physical hardware and logical workload management.

1. Real-time physical telemetry. Having real-time visibility into electrical usage, ambient conditions, and liquid cooling telemetry directly within the same console used to orchestrate bare metal server nodes drastically reduces mean time to resolution (MTTR) during hardware incidents.

On a GB300 NVL72 rack running near its ~150 kW ceiling, a supply-temperature drift or a stalled CDU pump needs to surface in seconds, not after a GPU has already throttled or faulted.

2. Precise physical space and capacity planning. Managing available rack units in real time ensures operators know exact spatial and power capacities before provisioning new bare-metal nodes or scheduling hardware maintenance. For example, this could be a 42U rack of conventional servers or a purpose-built 48U-class cabinet housing a GB200 NVL72 or GB300 NVL72 system.

At 120–150 kW per rack, power headroom is often the binding constraint well before rack-unit space runs out.

3. Integrated power controls. Providing low-level out-of-band (OOB) and IPMI power controls within a secure, centralized console allows remote engineering teams to manage individual nodes, GPU trays, or entire NVL72-class racks safely across dispersed data center locations.


Operational Impact

For neocloud operators, bringing physical rack topology, environmental telemetry, and hardware controls into a unified management plane yields immediate dividends:

  1. Minimized downtime. Instant leak detection and automated thermal alerts prevent catastrophic hardware failure before fluid damages dense GPU nodes. This kind of failure is extremely expensive on a $3M+ GB300 NVL72 rack than on a legacy air-cooled server.
  2. Accelerated hardware provisioning. Visualizing rack capacity and node placement speeds up physical racking, cable mapping, and node onboarding workflows, whether staging a single GPU server or commissioning an entire GB200 NVL72 pod.
  3. Simplified operations. Eliminates the need for operational personnel to context-switch between legacy DCIM tools, IPMI consoles, and Kubernetes/bare-metal orchestrators.

As AI workloads continue to push data center power and thermal densities to unprecedented levels, unifying bare-metal telemetry with cloud orchestration software will remain a core competitive advantage for next-generation cloud providers.s