Happy Sunday. If you have been following my pieces on data-center cooling, the fiber-optic supply chain, AI capital spending, and inference, you have probably noticed the pattern: I keep returning to the parts of the data center that sit behind the GPU. Networking belongs near the top of that list. It receives a fraction of the attention even though a large cluster is only as useful as its ability to move data between accelerators.
Most discussions of AI infrastructure start with compute. The network determines how much of that compute can be used at once. It moves data among GPUs, servers, and storage while keeping latency low enough that expensive accelerators do not sit idle. It also sits in the middle of export controls, antitrust review, and the U.S.-China technology dispute. That makes Mellanox one of the most useful stories in the sector.
NVIDIA announced its $6.9 billion acquisition of Mellanox in March 2019. The target was an Israeli networking company, and the deal needed approval in the United States, Europe, and China. It closed in April 2020, before the generative-AI boom made high-speed interconnects a mainstream investor topic. NVIDIA had bought the part of the system it would soon need most.
I also built an interactive timeline of Mellanox and the acquisition.
Mellanox began in Yokneam, Israel, in 1999. Eyal Waldman founded it with Michael Kagan, Shai Cohen, and several other engineers who had worked together at Intel and Galileo Technology. They were starting the company as the dot-com boom drove a rapid increase in servers and exposed a less glamorous problem: the machines still needed a fast way to talk to one another. The name Mellanox was meant to evoke a smooth flow of data.
The team had known one another since their years at Intel Israel. Mellanox focused on chips, adapters, switches, and software for high-performance interconnects. InfiniBand was emerging as an industry standard for moving data among servers and storage with high bandwidth and low latency. It came from a broader industry effort, and Mellanox became one of its most important vendors.
Mellanox was fabless. Its engineers designed the silicon and outsourced manufacturing to foundries such as TSMC. One of the company’s technical strengths was RDMA, or Remote Direct Memory Access. RDMA lets one system move data directly into another system’s memory with little CPU involvement. For high-performance computing, that saved time and freed the processor for other work.
Fiber entered wherever distance and bandwidth made copper less practical. Mellanox sold adapters, switches, cables, and transceivers for both electrical and optical links. Multimode fiber handled many shorter data-center runs, while single-mode fiber served longer reaches. Mellanox supported both copper and fiber, choosing the medium by reach, bandwidth, and cost. That breadth gave the company control over more of the connection.
By the time Mellanox went public in 2007, it had established itself in high-performance computing. Its ConnectX adapters and later SwitchX products supported InfiniBand, Ethernet, or both, depending on the model. LinkX cables and transceivers filled in the physical layer. The combination of silicon, systems, and software helped Mellanox win business with supercomputer operators and large cloud providers. Years before generative AI, it had built many of the tools needed to connect large accelerator clusters.
NVIDIA and Mellanox had worked together for years, pairing GPUs with high-speed networking in supercomputers and early AI systems. By 2018, the logic of owning both sides had become clearer. Faster GPUs created more pressure on the network, and NVIDIA had limited control over that part of the system. It agreed to pay $6.9 billion for Mellanox in March 2019, including a meaningful premium. The price looked full at the time. In hindsight, it was one of NVIDIA’s better strategic deals.
The transaction needed antitrust approval across several jurisdictions. U.S. and European reviews mattered, but China was the final major hurdle. Beijing was examining the deal in the middle of a worsening U.S.-China technology dispute, and approval took more than a year. The delay turned a networking acquisition into a geopolitical test.
In April 2020, China’s State Administration for Market Regulation approved the transaction with conditions. NVIDIA and Mellanox had to continue supplying Chinese customers fairly, avoid discriminatory bundling, and preserve interoperability. NVIDIA closed the acquisition soon afterward.
Those conditions did not disappear after closing. By 2024, Chinese regulators had opened an antitrust investigation into NVIDIA, including whether the company had honored the Mellanox commitments. As of August 2025, that investigation remained open. The original deal therefore connected three technology centers: Israeli engineering, an American buyer, and a Chinese regulator with leverage over the closing.
The timing helped: the acquisition closed in April 2020, about two and a half years before ChatGPT pushed generative AI into the mainstream. NVIDIA could not have predicted that exact moment. Jensen Huang had already been arguing that data centers would move toward accelerated computing, however, and Mellanox filled a real gap in that thesis. A GPU cluster needed a network designed for collective work at scale.
CUDA was already NVIDIA’s software advantage, and Mellanox gave the company a systems advantage to match. InfiniBand, Ethernet, RDMA, network adapters, switches, cables, and DPUs gave NVIDIA more control over how a cluster behaved after the GPUs were installed. The acquisition brought networking inside NVIDIA’s product roadmap instead of leaving it as outside plumbing.
The Mellanox portfolio can be understood through three product families: InfiniBand, RoCE-capable Ethernet, and BlueField DPUs. Fiber optics supports each one where reach or bandwidth requires it, although the portfolio’s value extends beyond the cable. Much of that value sits in congestion control, data movement, switching, and software.
Ethernet is the default network for most data centers because it is widely supported, flexible, and familiar to operators. InfiniBand is a separate fabric designed for low latency, high throughput, and predictable behavior in tightly managed clusters. Both can connect servers, storage, and accelerators. The choice usually comes down to the control and performance of a specialized fabric versus the economics and openness of Ethernet.
InfiniBand was Mellanox’s original center of gravity. RDMA lets data move between system memories with little CPU intervention, reducing overhead on communication-heavy workloads. NVIDIA’s Quantum switches and ConnectX adapters combine that feature with congestion control and network-management software. Copper can serve short links, while optical transceivers carry higher-speed traffic over longer distances. For large training clusters, that end-to-end control is the appeal.
Before the acquisition, InfiniBand was best known in supercomputing. NVIDIA pushed it deeper into AI training, where thousands of accelerators repeatedly exchange gradients and parameters. The fabric still requires careful design and operations, especially as clusters grow. Features such as SHARP move parts of collective communication into the network and reduce the amount of work the GPUs and hosts must handle.
RoCE brings RDMA semantics to Ethernet. It lets operators use the broader Ethernet ecosystem while retaining direct memory access and lower CPU overhead. Mellanox’s ConnectX adapters supported high-speed Ethernet as well as InfiniBand, depending on the product. For cloud operators, that offered a practical way to build AI networks while keeping much of the existing Ethernet stack.
RoCE’s flexibility comes with engineering work. Packet loss, congestion, and traffic bursts have to be managed carefully if the network is going to behave like a high-performance fabric. NVIDIA continued to improve the adapters, switches, telemetry, and congestion-control software after the acquisition. ConnectX-8, introduced in 2025, extended the portfolio to the next generation of AI systems, particularly Ethernet-based clusters.
BlueField added another layer. These programmable data-processing units could handle networking, storage, and security tasks that would otherwise consume CPU cycles. That left more of the host processor available for applications and gave operators a place to enforce policies close to the network. High-speed SerDes and optical ports connected the DPU to the rest of the system.
Mellanox had been developing this offload model before the acquisition. NVIDIA tied BlueField more closely to its software and systems roadmap afterward. By 2025, the same integration logic was reaching the optical layer through work on co-packaged optics. Moving optical components closer to the switch silicon promised lower power and higher density, although the technology still faced packaging, serviceability, and ecosystem challenges.
NVIDIA’s sales pitch changed after the deal. CUDA already gave developers a reason to choose its GPUs; Mellanox gave data-center operators a way to connect thousands of them. NVLink handled much of the scale-up traffic within a rack or system, and InfiniBand or Ethernet carried scale-out traffic across the cluster. NVIDIA could now sell the architecture around the accelerators as well as the accelerators themselves.
Scaling training clusters: InfiniBand, ConnectX, and the surrounding software let operators add GPUs while keeping communication overhead under control. The payoff showed up in cluster utilization and training time, with results varying by workload.
Serving inference efficiently: RoCE and high-speed Ethernet gave cloud operators a familiar base for workloads with different traffic patterns and cost targets. The network could be tuned for latency, throughput, and power.
Building a second growth engine: Mellanox turned networking into a multi-billion-dollar quarterly business for NVIDIA. The value of the acquisition was visible in revenue and in the company’s ability to sell complete systems.
Controlling more of the stack: NVIDIA could coordinate GPUs, NICs, switches, cables, and software. Customers still had architectural choices, while NVIDIA owned more of the performance envelope and captured more of the spending.
GPUs remain the headline product, and CUDA remains the software moat investors understand best. Mellanox explains how NVIDIA turned those GPUs into clusters. The deal brought a deep networking portfolio into the company just before AI infrastructure spending accelerated. It also put NVIDIA in the middle of a complicated supply-chain and regulatory story spanning Israel, the United States, and China. That combination is why I still think the Mellanox acquisition is one of the defining deals of the GPU era.







