Hello,
I am trying to get ET setup on the Quartz cluster at Indiana University (https://kb.iu.edu/d/qrtz). Simfactory compiles, however, I've run into a problem when a simulation begins and keep getting a segmentation fault (one per MPI process). The cluster uses SLURM, which is new to me, and RHEL 8. I've attached the relevant portion of an example output file (with stderr and stdout merged), the corresponding batch file that produced it when submitted, and below are the key settings from the machine file and the one line from the generic mpi runscript that I have edited. Any guidance would be much appreciated.
INI settings: max-num-smt = 1 # no hyperthreading on system num-smt = 1 ppn = 128 # physical cores mpn = 2 # NUMA domains ("sockets") max-num-threads = 128 # do not oversubscribe PUs num-threads = 4 # using 4 threads per process; make 2 may be better nodes = 1 # use more for actual runs
Edited runscript line: srun --nodes=@NODES@ --ntasks-per-node=@NODE_PROCS@ --cpus-per-task=@NUM_THREADS@ @EXECUTABLE@ -L 3 @PARFILE@
Thank you, Jessica
Dr. Jessica S. Warren Physics Lecturer Indiana University Northwest warrenjs@iun.edu
Hello Jessica,
Thank you for the complete log files attached.
I am trying to get ET setup on the Quartz cluster at Indiana University (https://urldefense.com/v3/__https://kb.iu.edu/d/qrtz__;!!DZ3fjg!6NAU24hNx0p3... ). Simfactory compiles, however, I've run into a problem when a simulation begins and keep getting a segmentation fault (one per MPI process). The cluster uses SLURM, which is new to me, and RHEL 8. I've attached the relevant portion of an example output file (with stderr and stdout merged), the corresponding batch file that produced it when submitted, and below are the key settings from the machine file and the one line from the generic mpi runscript that I have edited. Any guidance would be much appreciated.
My guess would be that there is a mismatch between the MPI library used to compile and the one used to run.
You can look at the file
configs/sim/bindings/Configuration/Capabilities/make.MPI.defn
which will list explicitly what libraries are used to link against and what directories were used. These should match the runtime directories that show up in the error message:
/N/soft/rhel8/openmpi/gnu/4.0.5
Sometimes one cannot really compile on compute nodes, did you check with the cluster admins whether they suppor this or whether compilation should only be done on the login nodes?
It is possible that "mpirun" used by the run script is not the one matching the MPI library used also there is the possibility that srun does not use the correct MPI stack. One would hope that "module load openmpi" would take care of this.
Finally it will be very helpful to first try a simple Hello, world MPI code such as this one (say):
https://mpitutorial.com/tutorials/mpi-hello-world/
and following one of the cluster examples.
To mimic how Cactus compiles you would not use the mpicc compiler wrapper but instead do something like (after inspecting make.MPI.defn):
gcc -L/N/soft/rhel8/openmpi/gnu/4.0.5/lib \ -I/N/soft/rhel8/openmpi/gnu/4.0.5/include \ -Wl,--rpath,/N/soft/rhel8/openmpi/gnu/4.0.5/lib \ hello.c -lmpi_cxx -lmpi -o hello
then submit a job for hello.
Yours, Roland
Hello Jessica,
You may also find something useful in the setting up a new machine seminar presentation:
https://www.einsteintoolkit.org/seminars/2022_02_24/index.html
Yours, Roland
Hi Roland,
Thank you so much. The compute nodes are able to be used for compilation, and the directories match what is listed in make.MPI.defn. When doing the 'hello' example you linked to, it was unable to compile due to a linker error (/usr/bin/ld: cannot find -lmpi_cxx). I re-ran it in verbose mode and found the directory it was searching did exist and did have lmpi but not lmpi_cxx. The admins said they had had some issues installing openmpi (couldn't recall exactly what), and recommended mpavich (since that does have lmpicxx installed and is their preferred implementation). However, they reinstalled openmpi in an effort to get that to work and it did allow the 'hello' script to compile, but when executed it produced:
-------------------------------------------------------------------------- No OpenFabrics connection schemes reported that they were able to be used on a specific port. As such, the openib BTL (OpenFabrics support) will be disabled for this port.
Local host: h1 Local device: mlx5_0 Local port: 1 CPCs attempted: rdmacm, udcm -------------------------------------------------------------------------- Hello world from processor h1.quartz.uits.iu.edu, rank 0 out of 1 processors
Similarly, doing the TOV job via sbatch, after the srun command it gave the same OpenFabrics message (for each MPI rank) and then the same segmentation faults as before. I've contacted the admins about this and am waiting to hear back. Do you have any recommendations - perhaps it would be easier to try switching over to mvapich? If so, could you point me to some resources on how to reconfigure?
Thank you, Jessica
Dr. Jessica S. Warren Physics Lecturer Indiana University Northwest warrenjs@iun.edu ________________________________ From: Roland Haas rhaas@illinois.edu Sent: Tuesday, August 9, 2022 9:48 AM To: Warren, Jessica Sawyer warrenjs@iun.edu Cc: users@einsteintoolkit.org users@einsteintoolkit.org Subject: [External] Re: [Users] Running with SLURM
Hello Jessica,
You may also find something useful in the setting up a new machine seminar presentation:
https://www.einsteintoolkit.org/seminars/2022_02_24/index.html
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://pgp.mit.edu .
Hello Jessica,
If you get the same error from hello-world and from Cactus then it would seem that there is still something off with the MPI stack.
The -lmpi_cxx option instructs the linker to link in C++ bindings for MPI though for just the hello world example, it being C code, this is not required and -lmpi alone is sufficient.
I would see two options that would let you get running somewhat quickly:
1. report your issues with OpenMPI and hello-world (including link to the source code on the web, and the exact command line to compile) to the admins and ask them for help
1.5 instead of using gcc to compile for OpenMPI do use the MPI official compiler wrapper mpicc which would just be:
mpicc -o hello hello.c
that is you do not have to pass and library or inlcude options. If this fails, I would definitely talk to the admins.
2. compile hello-world using mvapich. For this the easiest way is to make sure to load the mvapich module and then use the same compiler wrapper invication to compile:
mpicc -o hello hello.c
If 2 works then you can also compile the Einstein Toolkit with mvapich. You have to make sure to load the correct module before compiling the toolkit and then ExternalLibraries/MPI should figure out (from the mpicc wrapper) how to compile the toolkit.
Yours, Roland
Hi Roland,
Thank you so much. The compute nodes are able to be used for compilation, and the directories match what is listed in make.MPI.defn. When doing the 'hello' example you linked to, it was unable to compile due to a linker error (/usr/bin/ld: cannot find -lmpi_cxx). I re-ran it in verbose mode and found the directory it was searching did exist and did have lmpi but not lmpi_cxx. The admins said they had had some issues installing openmpi (couldn't recall exactly what), and recommended mpavich (since that does have lmpicxx installed and is their preferred implementation). However, they reinstalled openmpi in an effort to get that to work and it did allow the 'hello' script to compile, but when executed it produced:
No OpenFabrics connection schemes reported that they were able to be used on a specific port. As such, the openib BTL (OpenFabrics support) will be disabled for this port.
Local host: h1 Local device: mlx5_0 Local port: 1 CPCs attempted: rdmacm, udcm
Hello world from processor h1.quartz.uits.iu.edu, rank 0 out of 1 processors
Similarly, doing the TOV job via sbatch, after the srun command it gave the same OpenFabrics message (for each MPI rank) and then the same segmentation faults as before. I've contacted the admins about this and am waiting to hear back. Do you have any recommendations - perhaps it would be easier to try switching over to mvapich? If so, could you point me to some resources on how to reconfigure?
Thank you, Jessica
Dr. Jessica S. Warren Physics Lecturer Indiana University Northwest warrenjs@iun.edu ________________________________ From: Roland Haas rhaas@illinois.edu Sent: Tuesday, August 9, 2022 9:48 AM To: Warren, Jessica Sawyer warrenjs@iun.edu Cc: users@einsteintoolkit.org users@einsteintoolkit.org Subject: [External] Re: [Users] Running with SLURM
Hello Jessica,
You may also find something useful in the setting up a new machine seminar presentation:
https://urldefense.com/v3/__https://www.einsteintoolkit.org/seminars/2022_02...
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from https://urldefense.com/v3/__http://pgp.mit.edu__;!!DZ3fjg!9JAgxc4juluJwklwTQ... .
users@lists.einsteintoolkit.org