
Date: Friday, September 11, 2026
Location: Computing and Information Science Building room 350, Ithaca Campus
Click here to attend via Zoom
Speaker: Guy Amir, computer science postdoctoral researcher, The University of Texas at Austin
Title: Verifiable Fault Tolerance in AI Training
Abstract: Modern AI training is distributed across many machines. This makes it possible to train large models efficiently, but also makes training vulnerable to GPU and network failures, which can interrupt execution and complicate recovery. Existing recovery mechanisms are often fairly ad hoc and, in particular, do not provide formal guarantees that recovery preserves the intended behavior of the training process. In this work, we explore a different approach: using program semantics to reason formally about recovery. We develop the theory behind this idea and show how it can be used to provide correctness guarantees for fault recovery in distributed AI training.