Activation checkpointing trades memory for compute by discarding intermediate activations during forward passes and recomputing them during backpropagation. Arsh Koneru, an ML infrastructure intern, was tasked with building an improved activation checkpointing scheme that minimizes extra compute while staying within a given memory budget, building on PyTorch's existing tools like torch.compile's AOT Autograd and min-cut optimization.