ReportGem ReportGem

Academic paper

Application Failures and Machine Computational Efficiency

Authors: Carlo Graziani, Bethany Lusch, and O. E. Bronson MesserPublished: 2026-08-05Paper ID: 2608.05408Category: cs.DCLicense: CC BY 4.0

Abstract

We present a framework for evaluating uptime efficiency of Exascale-class scientific computers when application failure rates are appreciable. This is the situation that confronts current leadership-class scientific computing platforms and large AI training installations. What distinguishes scientific computing platforms is the heterogeneity of their applications. We argue that this diversity requires that failure rates and mean intervals between failures should be specified in terms of \emph{usage} (e.g. node-hours) rather than time, as is currently customary. We consider the usage loss terms due to failures, to checkpointing, and to restart costs, and update the framework of Daly (2006) allowing users to specify optimal checkpointing usage intervals that minimize such losses. We derive the machine computational efficiency, which specifies the expected fractional resource allocation that is available for scientific computation. We illustrate the methodology using one year of production runtime data from the \emph{Frontier} supercomputer at Oak Ridge National Laboratory.

This public page contains bibliographic metadata and the author abstract. Use the reader for licensed document access.

Open licensed paper reader