Term

NVLink

別名: NVLink

Overview

最終更新: 2026年7月9日

NVLinkは、NVIDIAが開発したGPU同士を直接結合するための高速インターコネクト技術である。汎用バスであるPCI Expressに比べて高い帯域幅と低遅延を実現し、複数のGPUを一体的なメモリ空間として扱えるようにすることで、大規模な学習・推論ワークロードにおける並列処理効率を高める役割を担う。データセンター向けGPUプラットフォームの中核技術として、NVIDIAのAIインフラ戦略を支えている。

概要

NVLinkは単体のGPUカード内の接続にとどまらず、サーバ筐体やラック単位で多数のGPUを結合し、あたかも一つの巨大な計算資源として扱うための基盤技術として位置づけられる。GPU間の通信速度がボトルネックになりやすい大規模モデルの学習において、NVLinkによる高速な相互接続は演算資源全体の実効性能を左右する重要な要素となっている。

技術的位置づけ

NVIDIAは自社のGPUプラットフォームを単なる半導体製品としてではなく、インターコネクト、パッケージング、周辺CPUを含めたシステム全体として提供する戦略を取っており、NVLinkはその中心的な構成要素である。2026年4月には、次世代GPUプラットフォーム「Rubin」においてCoWoP(Chip-on-Wafer-on-PCB)と呼ばれる新たなパッケージング技術の採用が検討されていることが報じられ、GPUとインターコネクトを含む実装形態全体の刷新が進められていることがうかがえる。こうした動きは、NVLinkを含む相互接続技術がGPU単体の性能向上と並行して進化を続けていることを示している。

主要な動向

2026年5月には、NVIDIAがクラウドデータセンター事業者CoreWeaveに20億ドルの追加投資を行い、新CPU「Vera」の単体供給や5ギガワット規模のAIインフラ構築に向けた協業を進めていることが明らかになった。このような大規模クラスタ構築においては、GPUを相互に結合するNVLinkの役割が一段と重要になる。一方で2026年6月には、NVIDIAのFY2026第4四半期決算が、同社が半導体メーカーから次世代デジタル経済のインフラ企業へと変貌していることを裏付けたと報じられており、GPUとそれを束ねるインターコネクト技術を含めたエコシステム全体が競争優位の源泉となっている構図が浮き彻となった。

同時期には、業界全体でNVIDIA依存からの脱却を模索する動きも顕在化している。2026年5月にはMicrosoftが自社開発の次世代AIアクセラレータ「Azure Maia 200」を発表し、推論コストの削減とNVIDIA依存からの戦略的転換を図っていることが報じられた。また2026年6月には、SamsungとAMDがOpenAIを核にHBM4戦略で結託し、NVIDIAとSK hynix連合の牙城に挑む構図が報じられており、AI半導体市場全体の勢力図が変化しつつある。加えて2026年6月には、PCI-SIGが次世代インターコネクト規格「PCI Express 8.0」の仕様策定を開始したと発表し、2028年のリリースを目指して最大1TB/sの帯域幅実現や将来的な光コネクタへの移行可能性が示された。これはNVLinkのような専用インターコネクト技術と、汎用バス規格であるPCIeの双方が、AIワークロードの拡大に伴う帯域幅需要の高まりに応じて並行して進化していることを示す動きである。

さらに2026年6月には、世界最大規模のGPUクラスタ「Colossus」を保有するxAIが、保有する55万台のGPUのうちわずか11%しか実効的に活用できていないことが判明したと報じられた。ハードウェアの急速な拡張に対しソフトウェア側の整備が追いついていないことが要因とされ、GPUを結合するインターコネクトを含むシステム全体の運用効率が、単純な保有台数以上に重要であることを浮き彻にした事例といえる。

Mentioned Articles

17 件

Research Papers

5 件
  • Evaluating Modern GPU Interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect

    Ang Li, S. Song, Jieyang Chen, Jiajia Li, Xu Liu, Nathan R. Tallent, K. Barker

    2019300 件引用Semantic Scholar

    High performance multi-GPU computing becomes an inevitable trend due to the ever-increasing demand on computation capability in emerging domains such as deep learning, big data and planet-scale simulations. However, the lack of deep understanding on how modern GPUs can be connected and the real impact of state-of-the-art interconnect technology on multi-GPU application performance become a hurdle. In this paper, we fill the gap by conducting a thorough evaluation on five latest types of modern GPU interconnects: PCIe, NVLink-V1, NVLink-V2, NVLink-SLI and NVSwitch, from six high-end servers and HPC platforms: NVIDIA P100-DGX-1, V100-DGX-1, DGX-2, OLCF's SummitDev and Summit supercomputers, as well as an SLI-linked system with two NVIDIA Turing RTX-2080 GPUs. Based on the empirical evaluation, we have observed four new types of GPU communication network NUMA effects: three are triggered by NVLink's topology, connectivity and routing, while one is caused by PCIe chipset design issue. These observations indicate that, for an application running in a multi-GPU node, choosing the right GPU combination can impose considerable impact on GPU communication efficiency, as well as the application's overall performance. Our evaluation can be leveraged in building practical multi-GPU performance models, which are vital for GPU task allocation, scheduling and migration in a shared environment (e.g., AI cloud and HPC centers), as well as communication-oriented performance tuning.

  • 9.3 NVLink-C2C: A Coherent Off Package Chip-to-Chip Interconnect with 40Gbps/pin Single-ended Signaling

    Yingcan Wei, Y. C. Huang, Haiming Tang, N. Sankaran, Ish Chadha, D. Dai, Olakanmi Oluwole, V. Balan, Edward Lee

    202339 件引用Semantic Scholar

    NVLink-C2C is the enabler for Nvidia's Grace-Hopper and Grace Superchip systems, with 900GB/s link between Grace and Hopper, or between two Grace chips. The connection provides a unified, cache-coherent memory address space that combines system and HBM GPU memories for simplified programmability. This coherent, high-bandwidth, low-power, low latency connection between CPU and GPUs is key to accelerating the most complex AI and HPC workloads.

  • The Nvlink-Network Switch: Nvidia’s Switch Chip for High Communication-Bandwidth Superpods

    A. Ishii, Ryan Wells

    202230 件引用Semantic Scholar
  • Towards Memory Disaggregation via NVLink C2C: Benchmarking CPU-Requested GPU Memory Access

    Felix Werner, Marcel Weisgut, T. Rabl

    20259 件引用Semantic Scholar

    Memory disaggregation decouples compute and memory resources, enabling efficient use of resources. Several interconnect technologies provide cache-coherent access to remote memory regions, which eases the use of disaggregated memory. Recent NVIDIA-based systems use the NVLink C2C interconnect, which provides cache-coherent memory access between CPUs and GPUs and their memory. While GPUs and NVLink are widely used to accelerate complex workloads, NVLink’s viability for connecting memory-expansion devices to a CPU remains unexplored. In this work, we quantify the characteristics of NVIDIA’s Grace CPU when accessing GPU memory via NVLink to assess NVLink’s viability for memory expansion. We benchmark throughput and latency for memory accesses on an NVIDIA Grace-Hopper system. We evaluate memory expansion when the CPU accesses both CPU and GPU memory and quantify the performance of database index operations with data stored in GPU memory. Our experiments show a throughput of up to 168 GB/s and access latencies between about 800 ns and 1000 ns.

  • FlexLink: Boosting your NVLink Bandwidth by 27% without accuracy concern

    Ao Shen, Rui Zhang, Junping Zhao

    20251 件引用Semantic Scholar

    As large language models (LLMs) continue to scale, multi-node deployment has become a necessity. Consequently, communication has become a critical performance bottleneck. Current intra-node communication libraries, like NCCL, typically make use of a single interconnect such as NVLink. This approach creates performance ceilings, especially on hardware like the H800 GPU where the primary interconnect's bandwidth can become a bottleneck, and leaves other hardware resources like PCIe and Remote Direct Memory Access (RDMA)-capable Network Interface Cards (NICs) largely idle during intensive workloads. We propose FlexLink, the first collective communication framework to the best of our knowledge designed to systematically address this by aggregating these heterogeneous links-NVLink, PCIe, and RDMA NICs-into a single, high-performance communication fabric. FlexLink employs an effective two-stage adaptive load balancing strategy that dynamically partitions communication traffic across all available links, ensuring that faster interconnects are not throttled by slower ones. On an 8-GPU H800 server, our design improves the bandwidth of collective operators such as AllReduce and AllGather by up to 26% and 27% over the NCCL baseline, respectively. This gain is achieved by offloading 2-22% of the total communication traffic to the previously underutilized PCIe and RDMA NICs. FlexLink provides these improvements as a lossless, drop-in replacement compatible with the NCCL API, ensuring easy adoption.

よくある質問

NVLinkとは何ですか?
NVLinkは、NVIDIAが開発したGPU同士を直接接続する高速インターコネクト技術である。PCI Expressより高速・低遅延な通信を実現し、複数GPUを一体的なメモリ空間として扱う用途で用いられる。
NVLinkとPCI Expressはどう違うのですか?
PCI Expressは汎用的な拡張バス規格であり、NVLinkはNVIDIAが独自開発したGPU間専用の高速接続技術である。2026年6月にはPCIe 8.0の仕様策定が始まり、両技術は帯域幅需要の拡大に応じて並行して進化している。
NVLinkは今後どのように進化していくのですか?
2026年4月には、次世代GPUプラットフォーム「Rubin」でCoWoPと呼ばれる新パッケージング技術の採用が検討されていると報じられ、インターコネクトを含む実装形態全体の刷新が進行中である。
NVLinkはNVIDIAのデータセンター戦略にどう関わっていますか?
2026年5月にはCoreWeaveへの追加投資や5ギガワット級AIインフラの構築が進められており、大規模GPUクラスタを効率的に結合するNVLinkの役割が一層重要になっている。
NVLink依存から脱却する動きはありますか?
2026年5月にMicrosoftが自社製アクセラレータ「Azure Maia 200」を発表し、2026年6月にはSamsungとAMDがOpenAIと連携してHBM4戦略を進めるなど、NVIDIA陣営以外からの独自技術追求が広がっている。

External Mentions

10 件