
NVIDIA公式否定でも止まらぬ株安:Kyber延期報道でイビデンが10%安になった理由
SemiAnalysisが報じたNVIDIAの次世代ラック「Kyber NVL144」延期観測を受け、イビデンなどアジア基板株が急落した。NVIDIAは公式に否定しているが、一次情報は未検証のX投稿1本のみで、公式否定より延期報道の方が市場を強く動かした。
別名: NVLink
NVLinkは、NVIDIAが開発したGPU同士を直接結合するための高速インターコネクト技術である。汎用バスであるPCI Expressに比べて高い帯域幅と低遅延を実現し、複数のGPUを一体的なメモリ空間として扱えるようにすることで、大規模な学習・推論ワークロードにおける並列処理効率を高める役割を担う。データセンター向けGPUプラットフォームの中核技術として、NVIDIAのAIインフラ戦略を支えている。
NVLinkは単体のGPUカード内の接続にとどまらず、サーバ筐体やラック単位で多数のGPUを結合し、あたかも一つの巨大な計算資源として扱うための基盤技術として位置づけられる。GPU間の通信速度がボトルネックになりやすい大規模モデルの学習において、NVLinkによる高速な相互接続は演算資源全体の実効性能を左右する重要な要素となっている。
NVIDIAは自社のGPUプラットフォームを単なる半導体製品としてではなく、インターコネクト、パッケージング、周辺CPUを含めたシステム全体として提供する戦略を取っており、NVLinkはその中心的な構成要素である。2026年4月には、次世代GPUプラットフォーム「Rubin」においてCoWoP(Chip-on-Wafer-on-PCB)と呼ばれる新たなパッケージング技術の採用が検討されていることが報じられ、GPUとインターコネクトを含む実装形態全体の刷新が進められていることがうかがえる。こうした動きは、NVLinkを含む相互接続技術がGPU単体の性能向上と並行して進化を続けていることを示している。
2026年5月には、NVIDIAがクラウドデータセンター事業者CoreWeaveに20億ドルの追加投資を行い、新CPU「Vera」の単体供給や5ギガワット規模のAIインフラ構築に向けた協業を進めていることが明らかになった。このような大規模クラスタ構築においては、GPUを相互に結合するNVLinkの役割が一段と重要になる。一方で2026年6月には、NVIDIAのFY2026第4四半期決算が、同社が半導体メーカーから次世代デジタル経済のインフラ企業へと変貌していることを裏付けたと報じられており、GPUとそれを束ねるインターコネクト技術を含めたエコシステム全体が競争優位の源泉となっている構図が浮き彻となった。
同時期には、業界全体でNVIDIA依存からの脱却を模索する動きも顕在化している。2026年5月にはMicrosoftが自社開発の次世代AIアクセラレータ「Azure Maia 200」を発表し、推論コストの削減とNVIDIA依存からの戦略的転換を図っていることが報じられた。また2026年6月には、SamsungとAMDがOpenAIを核にHBM4戦略で結託し、NVIDIAとSK hynix連合の牙城に挑む構図が報じられており、AI半導体市場全体の勢力図が変化しつつある。加えて2026年6月には、PCI-SIGが次世代インターコネクト規格「PCI Express 8.0」の仕様策定を開始したと発表し、2028年のリリースを目指して最大1TB/sの帯域幅実現や将来的な光コネクタへの移行可能性が示された。これはNVLinkのような専用インターコネクト技術と、汎用バス規格であるPCIeの双方が、AIワークロードの拡大に伴う帯域幅需要の高まりに応じて並行して進化していることを示す動きである。
さらに2026年6月には、世界最大規模のGPUクラスタ「Colossus」を保有するxAIが、保有する55万台のGPUのうちわずか11%しか実効的に活用できていないことが判明したと報じられた。ハードウェアの急速な拡張に対しソフトウェア側の整備が追いついていないことが要因とされ、GPUを結合するインターコネクトを含むシステム全体の運用効率が、単純な保有台数以上に重要であることを浮き彻にした事例といえる。

SemiAnalysisが報じたNVIDIAの次世代ラック「Kyber NVL144」延期観測を受け、イビデンなどアジア基板株が急落した。NVIDIAは公式に否定しているが、一次情報は未検証のX投稿1本のみで、公式否定より延期報道の方が市場を強く動かした。

xAIは世界最大規模のAIクラスター「Colossus」を保有するが、その計算能力のわずか11%しか活用できておらず、新社長が2ヶ月以内に50%への改善を宣言した。これは、急速なハードウェア拡張に対しソフトウェア整備が追いつかず、MetaやGoogleに比べて実効的なGPU稼働率が著しく低いという構造的な課題を露呈している。

NVIDIAが発表した2026会計年度第4四半期(2025年11月〜2026年1月)決算は、同社がもはや単なる半導体メーカーではなく、次世代デジタル経済の基盤を完全に支配するインフラストラクチャー企業として君臨しているこ […]

NVIDIAによるAIインフラストラクチャへの支配力が、また新たな段階へと突入した。 2026年1月26日、NVIDIAはクラウドデータセンター事業者であるCoreWeaveに対し、新たに20億ドル(約3000億円規模) […]

生成AIブームが「実験」のフェーズから「実装と運用」のフェーズへと移行する中、Microsoftがシリコンレベルでの巨大な賭けに出た。2026年1月27日、同社は自社開発の次世代AIアクセラレータ「Azure Maia […]

生成AI革命の裏側には、華々しいモデルの性能向上とは対照的な、泥臭く、過酷なハードウェアの現実が存在する。NVIDIA H100をはじめとする最新鋭GPUは、驚異的な演算能力を持つ反面、その運用は極めて不安定だ。 サーバ […]

AI(人工知能)革命を支える半導体市場において、長らく続いたNVIDIAの絶対的な支配体制に、今、大きな地殻変動の兆しが見える。既にAI時代の寵児であるOpenAIが、NVIDIAへの過度な依存からの脱却を目指し、長年の […]

PCI-SIGは、次世代インターコネクト規格「PCI Express 8.0」の仕様策定を開始したと発表した。2028年の仕様リリースを目指し、生データレート256.0 GT/s(ギガトランスファー/秒)、x16構成で1 […]
AI半導体市場を牽引するNVIDIAが、次世代の「Rubin」GPUプラットフォームにおいて、革新的なパッケージング技術「CoWoP(Chip-on-Wafer-on-PCB)」の採用を本格的に検討していることが、Dig […]

米国の厳格な輸出規制という逆風の中、半導体大手NVIDIAが中国市場向けに新たなAIチップを投入する計画であると、Reutersが報じている。最新のBlackwellアーキテクチャをベースとしつつも、性能と機能を大幅に絞 […]

NVIDIAは今後4年間で米国の半導体サプライチェーンに数千億ドルを投資する計画を発表した。同社CEO Jensen Huang氏は、全体で約5,000億ドル相当の電子機器を調達予定であり、そのうち「数千億ドル」を米国内 […]

AMDやIntel、Meta、Microsoft、Googleなど9社のテクノロジー大手が、AI向けの新しい高速相互接続規格「Ultra Accelerator Link(UALink)」の標準化を目指すコンソーシアムを […]

MetaがLlama 3の大規模言語モデルのトレーニングを行う中で、NVIDIA H100 GPUの頻繁な故障に悩まされていたことが明らかになった。Metaが最近公開した研究によると、16,384基のNVIDIA H10 […]

AIアクセラレータ領域において、NVIDIAの支配は圧倒的だ。これに少しでも対抗しようと、先日Intel、Googleらは、CUDAプラットフォームからの脱却を目指したオープンソースのソフトウェア・スイートを開発する団体 […]

NVIDIAと言えば昔はゲーミングGPU、今はAI向けGPUでその名を轟かせているが、GPUのみならず、CPUとGPUを組み合わせた高性能コンピューティング(HPC)向けのスーパーコンチップも製造している。この、NVID […]

チップ設計の巨人であり現在はTenstorrentのCEOであるJim Keller氏は、NVIDIAが最近発表したBlackwell GPUアーキテクチャの研究開発費が100億ドルにも及んだことに対し、単に相互接続方式 […]

NVIDIAは、現行世代であるHopperアーキテクチャGPU「H100」と比較して最大5倍の性能向上を誇るという、次世代「Blackwell」アーキテクチャGPUと、それに基づくAIアクセラレータ「B200」GPUを正 […]
High performance multi-GPU computing becomes an inevitable trend due to the ever-increasing demand on computation capability in emerging domains such as deep learning, big data and planet-scale simulations. However, the lack of deep understanding on how modern GPUs can be connected and the real impact of state-of-the-art interconnect technology on multi-GPU application performance become a hurdle. In this paper, we fill the gap by conducting a thorough evaluation on five latest types of modern GPU interconnects: PCIe, NVLink-V1, NVLink-V2, NVLink-SLI and NVSwitch, from six high-end servers and HPC platforms: NVIDIA P100-DGX-1, V100-DGX-1, DGX-2, OLCF's SummitDev and Summit supercomputers, as well as an SLI-linked system with two NVIDIA Turing RTX-2080 GPUs. Based on the empirical evaluation, we have observed four new types of GPU communication network NUMA effects: three are triggered by NVLink's topology, connectivity and routing, while one is caused by PCIe chipset design issue. These observations indicate that, for an application running in a multi-GPU node, choosing the right GPU combination can impose considerable impact on GPU communication efficiency, as well as the application's overall performance. Our evaluation can be leveraged in building practical multi-GPU performance models, which are vital for GPU task allocation, scheduling and migration in a shared environment (e.g., AI cloud and HPC centers), as well as communication-oriented performance tuning.
NVLink-C2C is the enabler for Nvidia's Grace-Hopper and Grace Superchip systems, with 900GB/s link between Grace and Hopper, or between two Grace chips. The connection provides a unified, cache-coherent memory address space that combines system and HBM GPU memories for simplified programmability. This coherent, high-bandwidth, low-power, low latency connection between CPU and GPUs is key to accelerating the most complex AI and HPC workloads.
Memory disaggregation decouples compute and memory resources, enabling efficient use of resources. Several interconnect technologies provide cache-coherent access to remote memory regions, which eases the use of disaggregated memory. Recent NVIDIA-based systems use the NVLink C2C interconnect, which provides cache-coherent memory access between CPUs and GPUs and their memory. While GPUs and NVLink are widely used to accelerate complex workloads, NVLink’s viability for connecting memory-expansion devices to a CPU remains unexplored. In this work, we quantify the characteristics of NVIDIA’s Grace CPU when accessing GPU memory via NVLink to assess NVLink’s viability for memory expansion. We benchmark throughput and latency for memory accesses on an NVIDIA Grace-Hopper system. We evaluate memory expansion when the CPU accesses both CPU and GPU memory and quantify the performance of database index operations with data stored in GPU memory. Our experiments show a throughput of up to 168 GB/s and access latencies between about 800 ns and 1000 ns.
As large language models (LLMs) continue to scale, multi-node deployment has become a necessity. Consequently, communication has become a critical performance bottleneck. Current intra-node communication libraries, like NCCL, typically make use of a single interconnect such as NVLink. This approach creates performance ceilings, especially on hardware like the H800 GPU where the primary interconnect's bandwidth can become a bottleneck, and leaves other hardware resources like PCIe and Remote Direct Memory Access (RDMA)-capable Network Interface Cards (NICs) largely idle during intensive workloads. We propose FlexLink, the first collective communication framework to the best of our knowledge designed to systematically address this by aggregating these heterogeneous links-NVLink, PCIe, and RDMA NICs-into a single, high-performance communication fabric. FlexLink employs an effective two-stage adaptive load balancing strategy that dynamically partitions communication traffic across all available links, ensuring that faster interconnects are not throttled by slower ones. On an 8-GPU H800 server, our design improves the bandwidth of collective operators such as AllReduce and AllGather by up to 26% and 27% over the NCCL baseline, respectively. This gain is achieved by offloading 2-22% of the total communication traffic to the previously underutilized PCIe and RDMA NICs. FlexLink provides these improvements as a lossless, drop-in replacement compatible with the NCCL API, ensuring easy adoption.