NVIDIA H100NVL
NVIDIA Graphics Card
The NVIDIA H100 NVL is a Hopper-architecture data centre GPU designed specifically for large language model inference, carrying 94GB of HBM3 high-bandwidth memory per card with 3.9 TB/s of memory bandwidth. A built-in Transformer Engine with FP8 precision delivers several times the large-model inference throughput of the previous generation. The PCIe 5.0 dual-slot passively cooled design can be paired through an NVLink bridge for 600 GB/s of card-to-card bandwidth, and Multi-Instance GPU (MIG) allows partitioning into up to seven independent instances. SWIT Technology supplies the H100 NVL as part of complete 4U and 6U GPU server builds.
Specifications
| Architecture | NVIDIA Hopper architecture (GH100) 4th-generation Tensor Cores Built-in Transformer Engine with FP8 support |
|---|---|
| GPU Memory | 94GB HBM3 high-bandwidth memory 3.9 TB/s memory bandwidth |
| System Interface | PCIe 5.0 x16 |
| Interconnect | NVLink bridge support, 600 GB/s card-to-card bandwidth |
| Virtualisation | Multi-Instance GPU (MIG), up to 7 independent instances NVIDIA vGPU support |
| Power | 350W to 400W max power consumption (configurable) |
| Cooling & Form Factor | Passive cooling (chassis airflow) Dual-slot full-height full-length, for rack servers |
| Software Support | CUDA, TensorRT-LLM, NVIDIA AI Enterprise, NCCL |
| Use Cases | Large language model inference, generative AI services, deep learning training, HPC |