DATA CENTRE AI NETWORKING

How AMD Pensando AI Networking Hardware Solves GPU Stalls

TM
Techmediaglobal
| 4 min read
75%
RUNTIME SPENT NETWORKING
2.4Tbps
BANDWIDTH PER GPU
13%
FASTER JOB COMPLETION
33%
LOWER SWITCHING COST

Up to 75% of AI training runtime can be lost to networking rather than computation, according to research from Google and Microsoft. AMD's new Pensando Vulcano 800 AI NIC targets that bottleneck directly, promising faster job completion, greater fault tolerance and lower switching costs for large-scale AI clusters.

A Network Interface Card Built for GPU-Hungry Workloads

The Pensando Vulcano 800 AI NIC is designed to handle high-throughput AI training and distributed inference workloads. It slots into a standard rack and combines three 800Gbps NICs per GPU, delivering up to 2.4Tbps of total bandwidth to each GPU.

Since the network directly impacts compute utilisation, closing the gap between raw GPU power and actual usable throughput is central to getting value out of expensive AI infrastructure investments.

Balancing Scale-Out and Scale-Across Architecture

Raw bandwidth alone doesn't guarantee real-world cluster performance. To maximise efficiency, the Pensando NIC combines scale-out networking between nodes with scale-across networking between GPUs.

AMD Networking describes scale-out as the ability to distribute a workload across multiple nodes, expanding the effective GPU pool beyond what a single server can provide. For training, this means syncing workloads across hundreds of servers without slowdowns; for inference, it means handling heavy request traffic smoothly. Scale-across, meanwhile, extends performance beyond the rack to clusters that span multiple data centres, keeping bandwidth high and latency low over long distances.

Part of a Bigger Rack-Scale Vision: AMD Helios

The Pensando card is available as part of AMD Helios, a rack-scale AI infrastructure solution that brings AMD Instinct MI455X GPUs, 6th Gen AMD EPYC Venice CPUs, AMD Pensando networking and AMD ROCm software together into one open, co-designed platform.

Compute, memory, networking, software, power and cooling are designed as one system rather than separate parts bolted together, built for large-scale AI training and inference.

"Think of a data centre rack the way you would think of an engine: every part has to work with every other part, or none of it performs the way it should."

— Chitra Barvekar, Senior Silicon Design Engineer, AMD

Faster and More Resilient AI Clusters

AMD notes that throughput alone doesn't determine success under real-world conditions — job completion rates, fault tolerance and sustained daily use matter just as much. By eliminating retransmits, system stalls and recovery cycles, the Pensando NIC delivers up to 13% faster AI job completion times, based on AMD engineering silicon modelling and synthetic benchmark simulation.

The card also distributes traffic across independent data paths, resulting in up to 33% lower switching costs. If one link or component fails, system performance degrades gracefully rather than bringing an entire multi-million-dollar training job to a sudden halt. Paired with in-service diagnostics and rapid fault isolation, infrastructure teams can fix issues on the fly without scheduling disruptive, cluster-wide maintenance windows.

Key Takeaways

  • Up to 75% of AI training runtime can be spent on networking rather than compute, per Google and Microsoft research.
  • The Pensando Vulcano 800 AI NIC combines three 800Gbps NICs per GPU for up to 2.4Tbps of total bandwidth.
  • It blends scale-out networking between nodes with scale-across networking between GPUs and data centres.
  • AMD reports up to 13% faster AI job completion times and 33% lower switching costs.
  • Failures degrade performance gracefully instead of halting entire multi-million-dollar training jobs.
  • The NIC is part of AMD Helios, a co-designed rack-scale platform combining GPUs, CPUs, networking and software.
Tags: AI Networking Data Centre AMD GPU Infrastructure AI Training Rack-Scale Computing AMD Helios