Hello there,
Here is the error attached I encountered while evolving black holes binary. The work-around I guess is to increase the number of processes (if I am not wrong), but I am not quite sure how to do that. I need some help to get around this issue.
Thanks in advance
Best regards, KS
Hi Karima,
Yes, that parameter file requires at least 2 mpi processes to run.
You seem to have submitted using simfactory and ended up running on 1 mpi process awith 2 openmp threads. If you instead do something like
simfactory/bin/sim run bbhHr --cores=2 --num-threads=1
if you're on a workstation that doesn't use a queueing system or
simfactory/bin/sim submit bbhHr --cores=2 --num-threads=1
if you're on a cluster. This should give you 2 mpi processes running on 1 openmp thread each. You might have to add other options to simfactory if you did so in your original run.
Cheers,
Peter
On Monday 2021-01-18 11:41, KARIMA SHAHZAD wrote:
Date: Mon, 18 Jan 2021 11:41:04 From: KARIMA SHAHZAD 02141911015@student.qau.edu.pk To: "users@einsteintoolkit.org" users@einsteintoolkit.org Subject: [Users] TAT/ Slab Error
Hello there,
Here is the error attached I encountered while evolving black holes binary. The work-around I guess is to increase the number of processes (if I am not wrong), but I am not quite sure how to do that. I need some help to get around this issue.
Thanks in advance
Best regards, KS
Hello Karima,
Here is the error attached I encountered while evolving black holes binary. The work-around I guess is to increase the number of processes (if I am not wrong), but I am not quite sure how to do that. I need some help to get around this issue.
Assuming you use simulation factory you need to pass options to make sure it creates multiple MPI processes. The auto-detection code tends to choose a number of threads equal to the the number of cores on your laptop and and a single MPI rank.
Eg if you have a 4 core laptop then you have to use:
./simfactory/bin/sim submit testrun01 --cores 4 --num-threads 2 --parfile ...
which uses a total of 4 cores and starts 2 threads per MPI rank so that you end up with 4 / 2 = 2 MPI ranks each of which uses 2 threads.
Similarly on clusters, the basic idea is to set --num-threads so that there are more than 1 MPI rank started (ie --num-threads is larger than the value for --cores or --procs [which are synonyms for each other]).
If not using simulation factory you have to manually use mpirun and OMP_NUM_THREADS. Eg:
export OMP_NUM_THREADS=2
mpirun -np 2 /home/karima/simulations/bbhHr/SIMFACTORY/exe/cactus_sim -L 3 /home/karima/simulations/bbhHr/output-0000/BBHHigherRes.par
which starts 2 MPI ranks each will use 2 OpenMP threads.
Yours, Roland
Sorry, I just have a quick follow-up on this subject, to clarify. Am I misinterpreting when I read "num-threads is larger than the value for cores or procs"? The cluster I am using (ThornyFlat) gives me the formula: procs* num-stm = np*num-threads Disregarding simultaneous multithreading (num-stm = 1) and considering np>1, gives procs>num-threads What I heard at the last ETK meeting was that procs is nr. processes, and it will be divided by the num threads. I thought that if we divide procs by num-threads, we get np. Indeed, testing it, I get: --procs=24 --num-threads=1 <=> np=24 --procs =24 --num-threads=2 <=> np=12 _______________________ Maria C. Babiuc Hamilton, Ph.D. Professor, Department of Physics College of Science, Marshall University, 1 John Marshall Drive, Huntington, WV, 25755 Room S 257, Phone: (304)696-2754
________________________________ From: users-bounces@einsteintoolkit.org users-bounces@einsteintoolkit.org on behalf of Roland Haas rhaas@illinois.edu Sent: Monday, January 18, 2021 5:17 PM To: KARIMA SHAHZAD 02141911015@student.qau.edu.pk Cc: users@einsteintoolkit.org users@einsteintoolkit.org Subject: Re: [Users] TAT/ Slab Error
Hello Karima,
Here is the error attached I encountered while evolving black holes binary. The work-around I guess is to increase the number of processes (if I am not wrong), but I am not quite sure how to do that. I need some help to get around this issue.
Assuming you use simulation factory you need to pass options to make sure it creates multiple MPI processes. The auto-detection code tends to choose a number of threads equal to the the number of cores on your laptop and and a single MPI rank.
Eg if you have a 4 core laptop then you have to use:
./simfactory/bin/sim submit testrun01 --cores 4 --num-threads 2 --parfile ...
which uses a total of 4 cores and starts 2 threads per MPI rank so that you end up with 4 / 2 = 2 MPI ranks each of which uses 2 threads.
Similarly on clusters, the basic idea is to set --num-threads so that there are more than 1 MPI rank started (ie --num-threads is larger than the value for --cores or --procs [which are synonyms for each other]).
If not using simulation factory you have to manually use mpirun and OMP_NUM_THREADS. Eg:
export OMP_NUM_THREADS=2
mpirun -np 2 /home/karima/simulations/bbhHr/SIMFACTORY/exe/cactus_sim -L 3 /home/karima/simulations/bbhHr/output-0000/BBHHigherRes.par
which starts 2 MPI ranks each will use 2 OpenMP threads.
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from https://nam02.safelinks.protection.outlook.com/?url=http%3A%2F%2Fpgp.mit.edu... .
Hello Maria,
Am I misinterpreting when I read "num-threads is larger than the value for cores or procs"?
Yes you are indeed misreading. The example I gave was:
./simfactory/bin/sim submit testrun01 --cores 4 --num-threads 2
ie num-thread=2 and cores (which is a synomym for procs) is 4.
The cluster I am using (ThornyFlat) gives me the formula: procs* num-stm = np*num-threads Disregarding simultaneous multithreading (num-stm = 1) and considering np>1, gives procs>num-threads
What I heard at the last ETK meeting was that procs is nr. processes, and it will be divided by the num threads.
No, procs is not the number of processes. It comes from processORs from the time where "processor" and "processor core" was the same thing.
--procs is the number total number of threads to create. The ends up being the same as the number of (logical) cores to use, and (ignoring hyperthreading) the same as the number of physical cores.
See:
http://simfactory.org/info/documentation/userguide/processterminology.html
My advise is usually:
* forget about hyperthreading (at least until you have numbers that indicate to you that for your simulation target you actually see a benefit) * use --cores for the total number of physical cores you want * use --num-threads for the number of threads you want per MPI rank
everything else is complicated / error-prone / dangerous / mostly untested.
Yours, Roland
users@lists.einsteintoolkit.org