ReportGem ReportGem

Academic paper

On Topology's Role in ML Training Performance

Authors: Sarah McClure, Tegan Wilson, Brad Karp, Michael Mitzenmacher, Sylvia Ratnasamy, Scott Shenker, Minlan YuPublished: 2026-08-03Paper ID: 2608.01707Category: cs.NILicense: CC BY 4.0

Abstract

Modern machine learning training workloads run on large-scale networks of compute accelerators. The networks commonly deployed in these systems are typically variations of two basic topologies: the fat-tree Clos and the torus. In this paper, we derive analytical results the elucidate how the choice of topology shapes achievable performance for the small set of collective communication operations that underlies modern machine learning workloads. We also consider how these results change when we include additional factors such as network failures and job placement strategies. Overall, we find that one topology does not dominate in all cases, but that the Clos achieves better collective completion time in most cases and provides benefits in resilience and flexibility.

This public page contains bibliographic metadata and the author abstract. Use the reader for licensed document access.

Open licensed paper reader