Hi all,
Another question:
I'm trying to restart a simulation with a different number of processors than the original run. Is there something in particular I need to do to make it work? When I do, it gets stuck for hours reading the checkpoints. The checkpoint is distributed in a number of files corresponding to the original number of procs I used, should I recombine them in a particular way?
Thanks again!
--- *Dr. Luciano Combi* Postdoctoral Researcher Perimeter Institute for Theoretical Physics CITA National Fellow (U. of Guelph) ---
Hello Luciano ,
I'm trying to restart a simulation with a different number of processors than the original run. Is there something in particular I need to do to make it work? When I do, it gets stuck for hours reading the checkpoints. The checkpoint is distributed in a number of files corresponding to the original number of procs I used, should I recombine them in a particular way?
The issue is that when changing the number of MPI ranks the data needs to be reorganized and right now this means that each MPI rank will open every single file to look for data, which can overwhelm the file system.
A quick workaround is often to set:
CarpetIOHDF5::open_one_input_file_at_a_time = "yes"
which reduces IO contention.
If that is still too slow (this has happened only with many hundreds of MPI ranks though), then you can try the hacked version of CarpetIOHDF5 in the branch rhaas/map which contains an helper script that you can run offline to parse all information in the checkpoint files into a "map" file. At checkpoint recovery time the MPI ranks then read in the map file which tells them exactly where they need to look for their data, this significantly reduces IO issues. It is is not user friendly though and was an emergency hack and will most likely require some trial and error to get it right, so setting the parameter above would be my first attempt.
Yours, Roland
I see, thanks, Roland!
As a matter of fact, I had that option already activated, otherwise it would just give me a memory error.
I'm thinking of maybe restarting the simulation with openMP activated to speed up the process, do you think it will help? Otherwise, I will try your hack.
Cheers. Luciano
On Mon, Apr 15, 2024 at 10:19 AM Roland Haas rhaas@illinois.edu wrote:
Hello Luciano ,
I'm trying to restart a simulation with a different number of processors than the original run. Is there something in particular I need to do to make it work? When I do, it gets stuck for hours reading the checkpoints. The checkpoint is distributed in a number of files corresponding to the original number of procs I used, should I recombine them in a particular way?
The issue is that when changing the number of MPI ranks the data needs to be reorganized and right now this means that each MPI rank will open every single file to look for data, which can overwhelm the file system.
A quick workaround is often to set:
CarpetIOHDF5::open_one_input_file_at_a_time = "yes"
which reduces IO contention.
If that is still too slow (this has happened only with many hundreds of MPI ranks though), then you can try the hacked version of CarpetIOHDF5 in the branch rhaas/map which contains an helper script that you can run offline to parse all information in the checkpoint files into a "map" file. At checkpoint recovery time the MPI ranks then read in the map file which tells them exactly where they need to look for their data, this significantly reduces IO issues. It is is not user friendly though and was an emergency hack and will most likely require some trial and error to get it right, so setting the parameter above would be my first attempt.
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://pgp.mit.edu .
Hello Luciano,
As a matter of fact, I had that option already activated, otherwise it would just give me a memory error.
Hmm, ok.
I'm thinking of maybe restarting the simulation with openMP activated to speed up the process, do you think it will help? Otherwise, I will try your hack.
I would be surprised if OpenMP helped since this is all IO bound and there is no computation. Note that you must ensure that even if you do not use OpenMP the number of MPI ranks is the same as your final simulation, otherwise you will end up having to wait for the checkpoint recovery again.
Using my hack you will need to switch to branch rhaas/map:
git checkout rhaas/map
then recompile and make sure you also compile all the utilities (simfactory does that automatically, in Cactus itself this is make foo-utils).
This will give you a new utility called
hdf5_create_binary_map
and there is also a helper script hdf5_create_binary_map.sh (both should end up in exe/sim/*).
Run hdf5_create_binary_map.sh in each checkpoint file, it will produce one "map" file per checkpoint file. This can be done in parallel (one invocation per file, all invocations in parallel) if the cluster admins let you (might need a short term interactive job maybe).
Concatenate all map files to a new map file using the same basename:
cat foo.file_*.map >foo.map
and make sure the concatenated map is in the same location as the checkpoint files.
The logic for this is mostly in the ReadMap function of CarpetIOHDF5/src/Input.cc
You may want to add a
CCTK_VINFO("Reading map files %s", fn);
just before the fopen call in there if you are not sure you have everything set up correctly. Otherwise there is no (obvious) indication that the map file is used (it silently falls back to the old / slow method if the map file is missing).
Yours, Roland
users@lists.einsteintoolkit.org