Hello all,
When I submit my simulation %%bash # start simulation segment ./simfactory/bin/sim submit NH --cores=1 --ppn-used=8 --walltime=0:2:00 Also tried this %%bash
./simfactory/bin/sim submit NH --cores=2 --num-threads=1 --walltime=0:20:00 again it gives the same warning it gives the warning that Total number of threads and number of cores per node are inconsistent: procs=1, ppn-used=8 (procs must be an integer multiple of ppn-used) and after that when i run the parameter file it does not run completely and shows only half or more than half output. How can I resolve this issue so that I get the complete output.
Hello Nisa,
When I submit my simulation %%bash # start simulation segment ./simfactory/bin/sim submit NH --cores=1 --ppn-used=8 --walltime=0:2:00 Also tried this %%bash
./simfactory/bin/sim submit NH --cores=2 --num-threads=1 --walltime=0:20:00 again it gives the same warning it gives the warning that Total number of threads and number of cores per node are inconsistent: procs=1, ppn-used=8 (procs must be an integer multiple of ppn-used) and after that when i run the parameter file it does not run completely and shows only half or more than half output. How can I resolve this issue so that I get the complete output.
If there is output missing then most likely the job was killed by the queuing system since it ran out of walltime. Note that the first command requested only 2 minutes of walltime which is almost certainly too short for any "real" run.
Usually this will show up at the bottom of the *.err file.
You can either let simfactory print both the *.out and the *.err file to screen (or pipe into less) using:
./simfactory/bin/sim show-output NH | less
or query where the simulation output directory is:
./simfactory/bin/sim get-output-dir NH
then use cd to go there and less to take a look at the err file.
The other option is that the job hung, which will usually also show up as the queueing system killing your run due to it running out of walltime, but also will typically mean that the last output (timestamp of the output files eg *.asc visible via ls -l) is much older than the time the job was killed by the queuing system.
If there is no queueing system (laptop) then something else could kill the job (eg runs out of memory).
The warning about ppn-use is due to inconsistent options. Namely you are claiming via ppn-used=8 to use 8 cores per node but then are requesting only 1 core. It is just a warning though, if the job started then you do not have to worry. If you would like to avoid the warning you could use --cores 1 --ppn-used 1. Does his happen on a cluster (private? One officially supported by the ET?)? Or you laptop that you auto-configured via "sim setup-silent" or on the tutorial server?
Yours, Roland
This happens on the laptop that I autofigured by following the tutorial. After the warning the the job has been started I think the issue is that the memory has been killed and it stops further processing of the simulation.
On Wed, 30 Mar 2022, 4:35 am Roland Haas, rhaas@illinois.edu wrote:
Hello Nisa,
When I submit my simulation %%bash # start simulation segment ./simfactory/bin/sim submit NH --cores=1 --ppn-used=8 --walltime=0:2:00 Also tried this %%bash
./simfactory/bin/sim submit NH --cores=2 --num-threads=1
--walltime=0:20:00
again it gives the same warning it gives the warning that Total number of threads and number of cores per node are inconsistent: procs=1, ppn-used=8 (procs must be an integer multiple of ppn-used) and after that when i run the parameter file it does not run completely
and
shows only half or more than half output. How can I resolve this issue so that I get the complete output.
If there is output missing then most likely the job was killed by the queuing system since it ran out of walltime. Note that the first command requested only 2 minutes of walltime which is almost certainly too short for any "real" run.
Usually this will show up at the bottom of the *.err file.
You can either let simfactory print both the *.out and the *.err file to screen (or pipe into less) using:
./simfactory/bin/sim show-output NH | less
or query where the simulation output directory is:
./simfactory/bin/sim get-output-dir NH
then use cd to go there and less to take a look at the err file.
The other option is that the job hung, which will usually also show up as the queueing system killing your run due to it running out of walltime, but also will typically mean that the last output (timestamp of the output files eg *.asc visible via ls -l) is much older than the time the job was killed by the queuing system.
If there is no queueing system (laptop) then something else could kill the job (eg runs out of memory).
The warning about ppn-use is due to inconsistent options. Namely you are claiming via ppn-used=8 to use 8 cores per node but then are requesting only 1 core. It is just a warning though, if the job started then you do not have to worry. If you would like to avoid the warning you could use --cores 1 --ppn-used 1. Does his happen on a cluster (private? One officially supported by the ET?)? Or you laptop that you auto-configured via "sim setup-silent" or on the tutorial server?
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
Also I get the same issue when I run and submit the simulation on the official Einstein Toolkit servel by using the jupyter note book
On Wed, 30 Mar 2022, 6:44 am Nisa Amir, nisaamir@math.qau.edu.pk wrote:
This happens on the laptop that I autofigured by following the tutorial. After the warning the the job has been started I think the issue is that the memory has been killed and it stops further processing of the simulation.
On Wed, 30 Mar 2022, 4:35 am Roland Haas, rhaas@illinois.edu wrote:
Hello Nisa,
When I submit my simulation %%bash # start simulation segment ./simfactory/bin/sim submit NH --cores=1 --ppn-used=8 --walltime=0:2:00 Also tried this %%bash
./simfactory/bin/sim submit NH --cores=2 --num-threads=1
--walltime=0:20:00
again it gives the same warning it gives the warning that Total number of threads and number of cores
per
node are inconsistent: procs=1, ppn-used=8 (procs must be an integer multiple of ppn-used) and after that when i run the parameter file it does not run completely
and
shows only half or more than half output. How can I resolve this issue so that I get the complete output.
If there is output missing then most likely the job was killed by the queuing system since it ran out of walltime. Note that the first command requested only 2 minutes of walltime which is almost certainly too short for any "real" run.
Usually this will show up at the bottom of the *.err file.
You can either let simfactory print both the *.out and the *.err file to screen (or pipe into less) using:
./simfactory/bin/sim show-output NH | less
or query where the simulation output directory is:
./simfactory/bin/sim get-output-dir NH
then use cd to go there and less to take a look at the err file.
The other option is that the job hung, which will usually also show up as the queueing system killing your run due to it running out of walltime, but also will typically mean that the last output (timestamp of the output files eg *.asc visible via ls -l) is much older than the time the job was killed by the queuing system.
If there is no queueing system (laptop) then something else could kill the job (eg runs out of memory).
The warning about ppn-use is due to inconsistent options. Namely you are claiming via ppn-used=8 to use 8 cores per node but then are requesting only 1 core. It is just a warning though, if the job started then you do not have to worry. If you would like to avoid the warning you could use --cores 1 --ppn-used 1. Does his happen on a cluster (private? One officially supported by the ET?)? Or you laptop that you auto-configured via "sim setup-silent" or on the tutorial server?
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
Hello Nisa,
Also I get the same issue when I run and submit the simulation on the official Einstein Toolkit servel by using the jupyter note book
Are you running a parfile of your own or the tov example? The tutorial server does not have lots of memory available and using too high resolution you can easily run out of memory.
Yours, Roland
I am using the parfile available in the gallery for binary neutron star mergers nsnstohmns.
On Wed, 30 Mar 2022, 6:43 am Roland Haas, rhaas@illinois.edu wrote:
Hello Nisa,
Also I get the same issue when I run and submit the simulation on the official Einstein Toolkit servel by using the jupyter note book
Are you running a parfile of your own or the tov example? The tutorial server does not have lots of memory available and using too high resolution you can easily run out of memory.
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
Hello Nisa,
I am using the parfile available in the gallery for binary neutron star mergers nsnstohmns.
According to the gallery page, this requires 8.8 GB of RAM. The tutorial server only has 8GB of memory so this run is quite likely to fail (see the output of the "free -h" command in a %%bash cell).
On your laptop I do not know of course.
While the HMNS is small for a NSNS simulation (too small actually as the resolution is about a factor of 2 too coarse to be reasonable), it is unfortunately a bit too large for the shared tutorial server. You are probably best off looking for an allocation at a local compute center or a workstation at your institute.
Yours, Roland
Okay thanks for your guidance. Can we run it on our computer by using the terminal as there is more space on the laptop.
On Wed, 30 Mar 2022, 7:04 am Roland Haas, rhaas@illinois.edu wrote:
Hello Nisa,
I am using the parfile available in the gallery for binary neutron star mergers nsnstohmns.
According to the gallery page, this requires 8.8 GB of RAM. The tutorial server only has 8GB of memory so this run is quite likely to fail (see the output of the "free -h" command in a %%bash cell).
On your laptop I do not know of course.
While the HMNS is small for a NSNS simulation (too small actually as the resolution is about a factor of 2 too coarse to be reasonable), it is unfortunately a bit too large for the shared tutorial server. You are probably best off looking for an allocation at a local compute center or a workstation at your institute.
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
Hello Nisa,
Okay thanks for your guidance. Can we run it on our computer by using the terminal as there is more space on the laptop.
Yes this run will likely run on your laptop using a terminal.
You can copy and paste the commands from the Jupyter cells (leaving out the %%bash line) and they will work on your laptop as well.
Yours, Roland
When I run the commands on my laptop the simulation just killed by itself. I am attaching the err file. Please help me in this regard.
On Thu, Mar 31, 2022 at 5:52 AM Roland Haas rhaas@illinois.edu wrote:
Hello Nisa,
Okay thanks for your guidance. Can we run it on our computer by using the terminal as there is more space on the laptop.
Yes this run will likely run on your laptop using a terminal.
You can copy and paste the commands from the Jupyter cells (leaving out the %%bash line) and they will work on your laptop as well.
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
Hello Nisa,
When I run the commands on my laptop the simulation just killed by itself. I am attaching the err file. Please help me in this regard.
I took a look at your err file but cannot tell much unfortunately. Usually a simulations is "killed" by the OS if it runs out of memory.
There may (or may not) be extra information available in the out file (usually it would refer to "OOM killer" or so). It would be good if you could provide it. There's a help page about what to provide to allow people to respond meaningfully:
https://einsteintoolkit.org/support.html
Namely:
--8<-- If you have trouble with a simulation run foo, send the foo.out and foo.err logs and the used parameter file if possible.
If a lot of text is copy-pasted into the email's body and makes it difficult to read, add the relevant files as attachments to the mail. They will be accessible from the archive page of the mailing list and all the subscribing users will be able to access them on their own email clients. --8<--
How long does the simulation run (this can be seen in the out file)?
How much actual memory does your laptop have (and how much of it is free when you start the simulation)?
When I run the parfile on my workstation Cactus consumes about 5.4GB of RAM (the gallery website states 8.8GB) when starting (but could use more as the simulation proceeds, in particular once the stars are about to touch) but the exact number may well differ between OS and computers.
What OS are you running on your laptop? Linux (not WSL)? Linux-subsystem-for-Windows (WSL, which version?)?
You can also try and directly run the simulation circumventing all of simfactory to see as much of the output as possible.
Usually something like, starting in the main Cactus directory:
export OMP_NUM_THREADS=4 mpirun -n 1 exe/cactus_sim par/nsnstohmns.par 2>&1 | tee nsnstohmns.log
will do this. It expects the parameter file in par/nsnstohmns.par and will redirect both the err and out output to nsnstohmns.log (as well as too screen).
You can also try to run "top" in a second terminal to see how memory is used up while the simulation starts.
This will use 4 threads (OMP_NUM_THREADS=4) and 1 MPI rank (-n 1) same as using "--cores 4 --num-threads 4" with simfactory.
The other output files will be a in directory nsnstohmns (or so) in the main Cactus directory.
Yours, Roland
users@lists.einsteintoolkit.org