BEGIN:VCALENDAR
VERSION:2.0
METHOD:PUBLISH
CALSCALE:GREGORIAN
PRODID:-//WordPress - MECv6.2.5//EN
X-ORIGINAL-URL:https://www.rc.fas.harvard.edu/
X-WR-CALNAME:FAS Research Computing
X-WR-CALDESC:HARVARD UNIVERSITY | FACULTY OF ARTS &amp; SCIENCES
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-PUBLISHED-TTL:PT1H
X-MS-OLK-FORCEINSPECTOROPEN:TRUE
BEGIN:VEVENT
CLASS:PUBLIC
UID:MEC-5c0321b6b78eecdfcf72e6a44222fef9@rc.fas.harvard.edu
DTSTART:20261022T150000Z
DTEND:20261022T190000Z
DTSTAMP:20260810T142300Z
RDATE;VALUE=PERIOD:20220607T170000Z/20220607T183000Z
CREATED:20260810
LAST-MODIFIED:20260810
PRIORITY:5
TRANSP:OPAQUE
SUMMARY:Checkpointing and Restarting Computational Workflows Workshop
DESCRIPTION:NOTE: This is an in-person session on campus.\nCheckpointing and Restarting Computational Workflows (in-person)\nDescription: In-person training on checkpointing and restarting computational workflows on FASRC clusters. This workshop will introduce general checkpointing strategies and demonstrate practical techniques for Python, C/C++, and Fortran applications, as well as AI and machine-learning workflows using PyTorch, TensorFlow/Keras, and JAX.\nParticipants will learn how to make long-running jobs resilient to failures, Slurm time limits, preemption, and planned requeue events. Hands-on exercises will cover saving and restoring application state, writing requeue-aware Slurm jobs, handling termination signals, and creating robust checkpoints for machine-learning training.\nNote: This is an in-person session. It will not be available via Zoom and will not be recorded.\nLevel: Intermediate\nTime: 11am - 3pm\nLocation: TBD (Northwest Labs vicinity)\nPresenters: Plamen Krastev\nWho can attend this workshop: Anyone with a FASRC cluster account. An active cluster account is required to participate in the hands-on exercises.\nWhat you will learn:\n\n\nWhy checkpointing is important for long-running and failure-prone computational workloads\n\n\nGeneral checkpointing approaches, design patterns, and restart granularity\n\n\nHow to create checkpoint-and-restart workflows in Python, C/C++, and Fortran\n\n\nHow to write requeue-aware Slurm jobs on the Cannon cluster\n\n\nHow to handle Slurm signals and exit gracefully before a job reaches its time limit\n\n\nHow to checkpoint AI/ML workflows using PyTorch, TensorFlow/Keras, Lightning, and JAX\n\n\nHow to save and restore model parameters, optimizer state, training progress, and random-number-generator state\n\n\nStrategies for distributed and large-scale ML checkpointing, including rank-0 and sharded checkpoints\n\n\nBest practices for reliable, portable, and storage-efficient checkpoints\n\n\nPrerequisites:\n\n\nA FASRC cluster account. If you do not have an account, see\nRequest a FAS Research Computing Account well in advance. See prerequisite 2 and 3.\n\n\nPrevious experience submitting batch jobs on a FASRC cluster. \n\n\nBasic familiarity with Slurm job scripts.\n\n\nBasic programming experience in Python, C/C++, or Fortran. Participants interested in the AI/ML examples should have basic familiarity with at least one supported machine-learning framework.\n\n\nRegistration: * Registration Link Pending *\n \n
URL:https://www.rc.fas.harvard.edu/events/checkpointing-workshop/
CATEGORIES:Training
END:VEVENT
END:VCALENDAR
