Nvidia Buys Hugging Face for Over 11 Billion Euros
Eva Mickler
4 min read Nvidia is acquiring Hugging Face for around 11.1 billion euros; the contract was signed on ...
Many IT leaders are expanding GPU capacity yet still waiting for response times. Free compute slots and rising inter-token latency simply don’t align. Cisco explicitly classifies AI inference on July 23, 2026 as a networking issue: data movement has become the primary bottleneck in GPU performance. What started as a symptom now demands an architectural decision.
RelatedLocal AI: Governance Before Hardware Purchase / Why Your Cloud Bill Never Gets Smaller
What is an AI back-end fabric? The AI back-end fabric links two networks: a North-South request network routes queries to clusters and returns results. An East-West GPU fabric (NVLink, InfiniBand, or RDMA over Ethernet) interconnects accelerators. Collectives such as All-Reduce govern step time because every training or inference step waits on data exchange between GPUs.
For IT leadership, the diagnostic picture is clear. Inter-token latency rises even when compute capacity is free. East-West collectives like All-Reduce dominate step time. Cisco summarized the core on July 23, 2026: “data movement is now the primary bottleneck for GPU performance.”
Operationally, the symptom resembles a capacity issue. Teams order more GPUs yet still measure longer wait times per token or per training step. The diagnosis shifts from compute resources to the data path between nodes. Overlooking this leads to funding unused accelerators while leaving the bottleneck in the fabric design.
Agentic workloads intensify pressure on the North-South side. A Cisco study forecasts agentic AI could multiply enterprise traffic ninefold by 2035. More requests hit clusters already constrained by East-West bandwidth and latency. IT leaders therefore need measurement points on both paths and clear ownership for the fabric that dictates step time.
AI inference relies on two connected but distinct fabrics. The north-south request network delivers load and returns responses. The east-west GPU fabric keeps the collectives running. NVLink, InfiniBand, or RDMA over Ethernet handle the high-speed coupling there. Errors in separating the two layers lead to misguided investment priorities.
If the request network scales while the GPU fabric stalls, inference remains slow. If the GPU fabric is strong but the request network is overloaded, wait times at the entry point rise. The decision therefore begins with a load map: which share of step time is consumed by all-reduce and related collectives, and which by ingress and egress. Without this breakdown, budget discussions remain speculative.
For infrastructure managers, this leads to an organizational consequence. Network, platform, and AI operations need shared metrics. Inter-token latency, collective duration, and port utilization belong in the same reviews as GPU utilization. Only then can leadership target the point where the bottleneck truly breaks.
The Dell’Oro Group marks 2025 as a turning point: Ethernet surpassed InfiniBand in market adoption for AI backend networks. Two years earlier, InfiniBand still held roughly 80 percent. When Dell’Oro launched its AI backend coverage at the end of 2023, InfiniBand accounted for about 80 percent of the market. The shift in adoption is therefore recent and steep.
Port speeds are accelerating in parallel. 800-Gbps switch ports have, within three years of first shipments, already surpassed the 20-million mark. 400 Gbps took six to seven years to reach that milestone. According to Dell’Oro’s forecast, by 2025 the majority of switch ports in AI backend networks will operate at 800 Gbps, by 2027 at 1,600 Gbps, and by 2030 at 3,200 Gbps. These are projections, not actual figures.
Dell’Oro expects 2026 to be the first year with volume shipments of 1.6-Tbps switches. The ramp-up is projected to outpace 800 Gbps and exceed 5 million ports within one to two years. Again, these are forecasts. For budget planning, this means the next port generation is already on the horizon while 800G is only now reaching broad deployment.
Economically, the shift is substantial. Ethernet in AI backends could generate around €69 billion in switch revenue within five years. Anyone planning fabrics today is operating in a market that refreshes port density, optics, and operational models on short cycles. Multi-year generational gaps are no longer the norm.
On June 11, 2025, the Ultra Ethernet Consortium released UEC Specification 1.0, an Ethernet communications stack for AI and HPC. Spec 1.0 targets modern RDMA over Ethernet and IP, open interoperability, and scaling to millions of endpoints. For IT leaders, this matters because it charts a path to multi-vendor AI fabrics built on Ethernet.
In parallel, the IEEE is accelerating the physical layer. IEEE 802.3df-2024 for 800 GbE was completed in about 16 months-four months ahead of schedule. The IEEE Task Force P802.3dj, targeting 200, 400, 800 Gb/s, and 1.6 Tb/s based on at least 200 Gb/s signaling, remains on track for completion in 2026. Standards and delivery timelines are aligning more closely than in earlier Ethernet cycles.
A single report on July 20, 2026 announced the release of Ultra Ethernet Specification 1.0.3, dated July 16, 2026, adding 200 Gb/s per lane. This information comes from a single source and should be treated cautiously. It does not confirm a final market standard but signals movement within the spec stack.
InfiniBand remains a visible counterpoint. In practitioner discussions throughout 2025 and 2026, InfiniBand is still cited for lower latency and higher all-reduce bandwidth. Well-tuned RoCEv2 can come close, but often falls 15 to 25 percent short in collective bandwidth compared to InfiniBand. Choosing Ethernet means choosing broader operational scope, supply chains, and ecosystems-and accepting the need to bridge potential collective-performance gaps through architecture and tuning.
Technical decisions collide with hard operational numbers. 800G modules typically consume 14 to 24 watts depending on reach and DSP, while 400G modules often sit at 10 to 12 watts but are less efficient per transmitted bit. Higher port speeds don’t just shorten step-time-they reshape cooling, power budgets, and rack layouts.
For a retail benchmark: an NVIDIA-compatible 800G OSFP-SR8 module for 50 meters of multimode fiber costs around €762 at FS.com. This is not a binding procurement price, but it gives a sense of optics pricing. Multiply by the port density of an AI fabric and the optics bill can quickly eclipse the switch hardware itself.
Then there’s the north-south cost lever in the cloud. Google Cloud Data Transfer Out to Europe in the Premium tier is free for the first GiB per month and account. After that, tiers run roughly €0.10 per GiB up to 1,024 GiB, about €0.095 per GiB up to 10,240 GiB, and around €0.074 per GiB beyond that. AWS charges Data Transfer Out to the internet in most US regions after the first 100 GB of the Free Tier at roughly €0.078 per GB, with descending tiers. The AWS figure is a secondary-source benchmark.
IONOS, one of Europe’s leading cloud and hosting providers, migrated from InfiniBand to Ethernet with Enterprise SONiC, RoCEv2, and EVPN-VXLAN, scaling a 400G data-center fabric. The example shows a real-world Ethernet-based AI and HPC-like architecture in a European context. It doesn’t replace a load test in your own workload, but it does prove that the stack of open network OS, RoCEv2, and EVPN-VXLAN can be deployed at scale in production.
First, map the symptom to your operational reporting. Track inter-token latency, all-reduce duration, and free GPU time in parallel. If latency rises while compute capacity is still available, prioritize the fabric and optics over additional accelerators. This protects CapEx and shortens the time to tangible performance gains.
Second, keep north-south and east-west traffic strictly separated in your architecture. Route paths, egress costs, and security zones belong in one view; collective bandwidth, RDMA behavior, congestion control, and port generations belong in another. Mixed calculations create false precision and flawed tenders.
Third, consciously choose your transport and operations stack. InfiniBand remains defensible where collective bandwidth and latency are the hard constraints and the 15–25 % gap of RoCEv2 would break the business case. Ethernet with RoCEv2, open specifications, and growing port economics shines where scalability, supply availability, multi-vendor operations, and integration into existing data-center fabrics matter. The UEC Specification 1.0 and the IEEE port roadmap support this option, but they don’t replace your own tuning or proof of concept.
Fourth, calculate the total cost across port, module, wattage, and data transfer. 800G-and soon 1.6 Tbps-boost throughput while shifting power and optics budgets. Cloud egress tiers determine whether inference results stay in your own fabric or drain budgets elsewhere. Bringing these four lines together early reveals whether the network sets the tempo-and lets you intervene before the next GPU order merely compounds the bottleneck at a higher price.
When inter-token latency rises despite available compute capacity and East-West collectives like All-Reduce dominate step time, the bottleneck lies in the network. Cisco describes data movement in this context as the primary performance bottleneck for GPUs.
The north-south request network handles requests to and from the cluster, while the east-west GPU fabric links accelerators via NVLink, InfiniBand, or RDMA over Ethernet. Collective operations on the east-west path often determine step time.
By 2025, Ethernet surpassed InfiniBand in adoption for AI backend networks. 800-Gbps ports exceeded 20 million units within three years. Dell’Oro forecasts that by 2025, most AI backend switch ports will operate at 800 Gbps, rising to 1,600 Gbps by 2027 and 3,200 Gbps by 2030. 2026 is projected as the first year with volume shipments of 1.6-Tbps switches, scaling to over 5 million ports within one to two years.
Not universally. In 2025 and 2026 discussions, InfiniBand continues to offer lower latency and higher All-Reduce bandwidth. Well-tuned RoCEv2 can come close, but often lags behind InfiniBand by 15 to 25 percent in collective bandwidth. Ethernet excels in ecosystem breadth, availability, and integration; the collective performance gap must be measured against your specific workload.
Module power draw, optical transceiver pricing, and cloud egress fees. 800G modules typically consume 14 to 24 watts. As a retail benchmark, an 800G OSFP-SR8 module retails around €762. Google Cloud egress to Europe costs approximately €0.10 to €0.074 per GiB after the first free GiB, depending on tier. AWS egress rates start at around €0.078 per GB after the free tier. Pricing for AWS and retail modules are indicative benchmarks, not binding contract values.
Image source: AI-generated (August 2026)
Read more on Digital Chiefs
Digital ChiefsAmazon and Alphabet: Negative Cash Flow, Long-Term CommitmentsDigital ChiefsLocal AI: Governance Before Hardware PurchaseDigital ChiefsAI Regulation: Up to 3 Percent of Corporate RevenueMore from the MBF Media Network
cloudmagazinCloud Scarcity: What Q2 Means for Procurement Teams securitytodayDrones Over Critical Infrastructure: Perimeter Transforms into Cyber-OT mybusinessfutureQ2 2026 Financing Climate: Loans Tight, Capital Available