普通外文研报
The Cool Aisle by George Notter... 11% MFU and a Special Request
研报英文原文证据摘录
The Cool Aisle by George Notter... 11% MFU and a Special Request
May 24, 2026
how the models are designed. Nonetheless, we think it's a good view on how model performance is trending. Jensen
– at a presentation at Stanford last week – called the MFU metric "just simply wrong," and he would "rather be at low
MFU." The point he made was simply that he would rather overprovision hardware than not having enough compute.
Also, Jensen noted that any given time, constraints in compute, memory, or network could limit the performance of the
whole system. This aligns with how we view the importance of networking within the broader AI infrastructure stack
– if network performance is bottlenecking GPU performance, Cloud Providers, Model Builders, and Neoclouds need to
pay more attention towards optimizing it.
Who Fixes the MFU Problem? In our view, networking is a significant factor in improving MFUs. Increased network
efficiency allows more of the time to be spent on GPU-hours – rather than in the network, as described above. In this
context, these xAI Colossus clusters are built using NVIDIA's SpectrumX Ethernet switching fabric. It's an important
detail for the competitive narrative at Arista – we've long considered them to be the "best-of-breed" switching provider.
By no means is the 11% MFU directly contributable to a deficiency in SpectrumX's switching infrastructure at xAI –
as we noted above, there are challenges with scaling GPU clusters at the scale that xAI has done to this point (550k
GPUs). However, if Arista is able to demonstrate meaningfully higher utilization benefits, it's possible it could open
the door for them at Neocloud accounts that have previously been "all-in" on NVIDIA's full stack. Elsewhere, vendors
本摘录由系统从所标注的 PDF 证据页直接提取并保留英文原文,不做批量翻译;登录后在阅读器切换中文时才按需翻译。
打开研报阅读器