⚠️ Kraken accounts are not granted by default, and must be requested by each individual staff member.
You can find the information & steps to request an account in the first section below, aptly titled "Request Kraken Account".
💭 If you aren't sure whether you need a kraken account, please ask a colleague or supervisor.
📡 For instructions on connecting to servers via SSH, please see this page.
🐛 Slurm is our cluster management and scheduling system. We will focus on a few slurm commands to get you started once you have requested and received your account to log into Kraken.
Use the form to request an account for the Kraken cluster. You'll need to have your SSH key pair already generated.
When you are ready, click the button that says Fill out form.
Mac users: To easily upload the id_rsa.pub file (RSA public key) you can run the following commands in order, and then be able to drag and drop the file from Finder:
cd ~
cp .ssh/id_rsa.pub ~/Desktop
open ~/Desktop
WSL users: To easily upload the id_rsa.pub file, you can run the following commands in order from WSL2, and then be able to drag and drop the file from File Explorer:
cd ~/.ssh
explorer.exe .
To avoid performance issues on our login servers, we ask that any and all large file transfers between Kraken and any other system or storage location are executed only on lyra.dfci.harvard.edu.
Question: Can I perform file transfers from ada, noah, k1, or k2?
Answer: ❌ No, because those systems are not lyra.
Question: I am not confident on the command line with transferring data, and would prefer a graphical interface (GUI) to transfer with. What are my options?
Answer: There are a few tools that make SFTP much easier with the usage of a GUI. Among the most popular of these tools is FileZilla. You will need to configure it to use your Unix password (or SSH keys) and provide the right login name, server name (the correct one is lyra.dfci.harvard.edu !), and maintain a persistent connection. You can find our tailored overview of FileZilla on this knowledgebase page.
sbatch is used to submit a job script
srun is used to run commands directly on a work node
squeue is used to list all running and queued jobs
How to Submit a Job
The easiest way to test submitting a job is to use the simple submit script below. Same the file below to a file named myJob. To submit this script, type sbatch myJob. If you get an output file, you have successfully submitted your first job on the kraken cluster.
#!/bin/bash
#SBATCH --job-name=myJob # job name - can use -J jobname for short
#SBATCH --error=/fullpathtoerrorfile # error file
#SBATCH --output=/fullpathtooutfile # output file
/bin/hostname
How to Use srun
The example below shows how to execute a basic srun so that I can attain a work node to run my command or software
[user@kraken1 ~]$ srun –pty bash
[user@node02 ~]$ hostname
node02
How to Use srun to Run software that Has Graphics
For this to work, you would need to have an X server with X11 enabled on the machine you ssh into kraken. Mac users can use XQuartz and Windows Users can use Xming or Exceed as X Servers.
[user@kraken1 ~]$ srun –x11 -p interactive –pty bash
How to Check on Currently Running and Queued Jobs
[user@kraken1 ~]$ squeue
JOBID PARTITION NAME USER ST TIME NODES NODELIST(REASON)
23998 defq bash user1 R 2:09:50 1 node01
24003 defq bash user2 R 12:24 1 node02
WIP
List job ids
squeue
# This will give you a list of jobids to use when determining what to cancel.
Cancel a Job
# To cancel jobid 123 OR to cancel all elements from job array 123:
scancel 123
# To cancel jobid or job array 100-120:
scancel {100..120}
# To cancel array id 4 and 5 from job array 100:
scancel 100_4 100_5
# To cancel array id 4-10 from job array 200:
scancel 200_[4-10]
# To cancel all of your running jobs (where yourusername is replaced with a real username):
scancel -u yourusername –state=running
# To cancel all of your pending jobs (where yourusername is replaced with a real username):
scancel -u yourusername –state=pending
Pausing and Resuming Jobs
# Suspend running job with id 123:
suspend 123
# Resume job 123:
resume 123
WIP
Simple Submit Script
#!/bin/bash
#SBATCH --job-name=myJob # job name - can use -J jobname for short
#SBATCH --error=/fullpathtoerrorfile # error file
#SBATCH --output=/fullpathtooutfile # output file
/bin/hostname
HINT: Type sbatch scriptname to submit this script
Simple Submit Script with More Options
#!/bin/bash
#SBATCH --job-name=myJob
#SBATCH --mem=2G # total memory need
#SBATCH --time=5:00:00 # DD-HH:MM:SS max walltime(time limit) requested - can use -t 5:00:00 for short
#SBATCH --error=/fullpathtoerrorfile # error file
#SBATCH --output=/fullpathtooutfile # output file
#SBATCH --mail-type=END,FAIL # email notification when job ends/fails
#SBATCH --mail-user=username@jimmy.harvard.edu # email to notify
module purge # clear loaded modules
module load gcc/7.2.0 # load a module
RUNDIR=/fullpathtorundir # set a variable
cd $RUNDIR
/fullpathtocommand
Array Submit Script
#!/bin/bash
#SBATCH --job-name=myJob
#SBATCH --array=1-16
#SBATCH --time=02:00:00
#SBATCH --ntasks=1
#SBATCH --mem=2G
#SBATCH --error=/fullpathtoerrorfile
#SBATCH --output=/fullpathtooutfile
cd /tofolder
/fullpathtocommand $SLURM_ARRAY_TASK_ID
Threaded or Multi-Processor Script - to run on 1 node with multiple cores per task
#!/bin/bash
#SBATCH --job-name=myJob
#SBATCH --time=02:00:00
#SBATCH --mem=2G
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=2 # number of cores per task - can use -c 1 for short - default is 1
#SBATCH --ntasks-per-node=1 # number of tasks per node - can use -n 1 for short - default is 1
#SBATCH --error=/fullpathtoerrorfile
#SBATCH --output=/fullpathtooutfile
/fullpathtocommand
MPI Script - to run on multiple nodes with one CPU per task
#!/bin/bash
#SBATCH --job-name=myJob
#SBATCH --time=02:00:00
#SBATCH --mem=2G
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=2 # number of cores per task - can use -c 1 for short - default is 1
#SBATCH --ntasks-per-node=1 # number of tasks per node - can use -n 1 for short - default is 1
#SBATCH --error=/fullpathtoerrorfile
#SBATCH --output=/fullpathtooutfile
/fullpathtocommand
MPI and Threaded Script - to run on multiple nodes with more than one CPU per task
#!/bin/bash
#SBATCH --job-name=myJob
#SBATCH --time=02:00:00
#SBATCH --mem=2G
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=2 # number of cores per task - can use -c 1 for short - default is 1
#SBATCH --ntasks-per-node=1 # number of tasks per node - can use -n 1 for short - default is 1
#SBATCH --error=/fullpathtoerrorfile
#SBATCH --output=/fullpathtooutfile
/fullpathtocommand
8 tasks with 4 cores per task to be divided up between 2 nodes
4 tasks with 16 cores total to run on each node
2 tasks to run on each socket (2 sockets per node)
We use Modules to load software on Kraken. When a software is loaded, its path and library path are already taken care of. In the examples below, the $ character denotes the command line, so you would run whatever command is after that symbol on the same line.
View available software
$ module available
———————————————— /cm/local/modulefiles ————————————————
cluster-tools-dell/8.2 cm-setup/8.2 cmsh gcc/8.2.0 module-git openldap
cluster-tools/8.2 cm-upgrade/8.2 dot ipmitool/1.8.18 module-info python2
cm-scale/8.2 cmd freeipmi/1.6.2 lua/5.3.5 null shared
View currently loaded software
$ module list
Currently Loaded Modulefiles:
1) gcc/8.2.0 2) slurm/18.08.4
Load a module and confirm
$ module load lapack/gcc/64/3.8.0
$ module list
Currently Loaded Modulefiles:
1) gcc/8.2.0 2) slurm/18.08.4 3) lapack/gcc/64/3.8.0
# We see the third Modulefile is the one we just loaded.
Unload a module and confirm
$ module unload lapack/gcc/64/3.8.0
$ module list
Currently Loaded Modulefiles:
1) gcc/8.2.0 2) slurm/18.08.4
# We see the third Modulefile is missing, which was the one we just unloaded.
Monitoring Commands #With description after comment sign
squeue ### List all running and queued jobs
squeue -u <username> ### List all running and queued jobs under a <username>
scontrol show JobId=9999 ### Get detailed information on job 9999
sstat -j 10421 ### Get detailed resource consumption of a job.
### Get detailed resource consumption of a job for the specified parameters:
sstat –format=AveCPU,AvePages,AveRSS,AveVMSize,JobID -j 10421 –allsteps
Monitoring info for past jobs
### Get detailed information of a finished job:
sacct -j 9999
### Get detailed information of a finished job for the specified parameters:
sacct -j 9999 –format=JobID,JobName,Partition,QOS,Elapsed,Start,NodeList,State,ExitCode
### Get detailed information of all finished jobs under a user for the specified parameters:
sacct -u username –format=JobID,JobName,Partition,QOS,Elapsed,Start,NodeList,State,ExitCode
Monitoring work nodes
### List status of each node:
sinfo
### Get detailed information of each node:
scontrol show nodes
### View detailed job and node information for each node via X-windows:
sview
Each user is given a home directory (/homesN/username) that is mounted on all cluster nodes. It has a size limit of 50Gb. Please use it for basic login scripts and simple submit jobs. It is backed up daily, and backups are kept for 6 months.
Each user has write permissions to their appropriate lab share (/name-of-lab.) Lab shares are mounted on all cluster nodes and can also be mounted on desktops and laptops. Size limits depend on the particular lab, this is where you put your regular data and work files. It is backed up weekly, and backups are kept for 3 months.
Every lab has access to our high performance scratch space (/cluster/name-of-lab) Each user can create their own folder. This filesystem is managed by the GPFS parallel filesystem and is appropriate for data intensive jobs. It is mounted on all work nodes, but not on the head nodes. It is considered as a temporary storage. Files are not backed up and if the storage fills up we may delete any files, so once your analysis has been completed please move your files to your lab share.
Every work node has a small storage partition (approximately 100Gb) that is suitable for temp files (/tmp). This partition is not backed up and files can be deleted at any time. It is best not to use it since it is specific to the node and not shared across nodes. If your application is contained within the same node you can point TMPDIR to it.
Click sections above to expand them.