#

Cluster Computing

Jump To:  CANNONFASSE | KEMPNER

Description 

Researchers can take advantage of the scale of the FASRC cluster by setting up workflows to split many different tasks into large batches, which are scheduled across the cluster at the same time.  Most clusters are made from commercial hardware, so do note that the performance of an individual single-threaded computation job may not be any faster than a local new workstation/laptop with new CPU and flash (SSD) storage; cluster computing is all about scale.  Doing computations at scale allows a researcher to test many different variables at once, thereby shorter time to outcomes, and also provides the ability to ask larger, more complex problems (I.e. larger data sets, longer simulation time, more degrees of freedom, etc.) 

Please Note: The FASRC research clusters cannot host class accounts and academics due to FERPA and other requirements. Please work with University RC and ATG to set up your course on the Academic Cluster and Canvas. See: https://atg.fas.harvard.edu/ondemand

Key Features and Benefits  

Service Expectations and Limits: 

Cluster-computing is not intended to be a larger version of a single machine. Scaling computations to take advantage of a cluster may require changes to how the computation is configured and executed.  As cluster-computing is a shared resource, many factors affect performance including other computations on the same compute nodes competing for memory, storage, and network. In addition, due to the nature of research that is being performed, it isn’t always well understood by the end user how scaling out their computations on a cluster will differ from performing them on a single device in terms of resource requirements (cores, memory, time, data storage), as these are managed in a different way on a single device.  It is always best for researchers to test smaller data sets and batches of tasks before submitting hundreds or thousands. 

At FASRC, availability, uptime, and backup schedule is provided as best effort with staff that do not have rotating 24/7/365 shifts.   

Cannon

Queuing System: SLURM 

Security Levels: Level 1 (DSL1), Level 2 (DSL2)

Shared:  Resources are provided by FASRC and are distributed evenly amongst the various lab groups. 

Dedicated: Resources are provided by Centers, Departments, or PIs 

  • Access: Restricted PI groups from Center/Department/Lab. 
  • Scheduling: Fairshare, see https://www.rc.fas.harvard.edu/fairshare/ for more details. 
  • Queue/Partition name: pi_name or center name (e.g. huce) 
  • Features: time limits vary, non-exclusive nodes, multi-node parallel 
  • Cost: Billing is done at the school level so please consult your school administrator.  

Backfill: Users have the ability to utilize unused parts of the dedicated queues. 

  • Access: All PI groups. 
  • Scheduling: Fairshare, see https://www.rc.fas.harvard.edu/fairshare/ for more details. 
  • Queue/Partition name: serial_requeue, gpu_requeue 
  • Backfilled by Kempner nodes
  • Features: See our partitions table at Running Job/Partitions
  • Cost: Billing is done at the school level so please consult your school administrator.  

As is customary with high-performance computing clusters home directories are provided to each individual user and scratch/temp space is provided to each PI Group.

Home directories - Isilon

  • Features:  Regular performance, snapshot, DR copy, network attached to cluster (NFS), quota, SMB 
  • Mount point: /home/username 
  • Quota: 100 GB
  • Cost: Included as part of Cluster Computing

Scratch -  Vast

  • Features: High-performance, temporary, single copy, network attached to cluster, quota (files + size), 90 day retention policy. 
  • Mount point: /n/netscratch/pi_lab
  • Quota: 50 TB max per Lab 
  • Cost: Included as part of Cluster Computing

All Research Facilitation services pertaining to the use of the cluster are included in the cost.

Available to: 

All PIs from a supported Harvard School with an active FASRC account.  For more information see Account Requests 

FAS Secure Environment (FASSE)

Queuing System: SLURM

Security Levels: Level 3 (DSL3)

See: https://docs.rc.fas.harvard.edu/kb/fasse/

Shared:  Resources are provided by FASRC and are distributed evenly amongst the various lab groups. 

  • Access: All PI groups. 
  • Scheduling: Fairshare, see https://www.rc.fas.harvard.edu/fairshare/ for more details. 
  • Queue/Partition name: shared, general, bigmem, test, gpu_test
  • Features: See our partitions table at Running Job/Partitions
  • Due to segregated nature, does not backfill to other partitions/clusters
  • Cost: Billing is done per project at the school level.  See Billing FAQ for more details.

Kempner

Queuing System: SLURM 

Security Levels: Level 1 (DSL1), Level 2 (DSL2)

Shared:  Resources are provided by FASRC and are distributed amongst the various Kempner groups. 

  • Access: Kempner PI groups. 
  • Scheduling: Fairshare, see https://www.rc.fas.harvard.edu/fairshare/ for more details. 
  • Queue/Partition name: CPU and GPU-specific partitions as well as all Cannon partitions (except Kempner/HMS labs who can only use Kempner partitions). See our partitions table at Running Job/Partitions
  • Backfills idle cycles  into the Cannon cluster
  • Features: Exclusive and non-exclusive nodes, multi-node parallel 
    Please note that some resources such as GPU queues may have different limits.
  • Cost: Billing is done at the school level.