Hello, In order to avoid stressing the filesystem on the cluster I'm running on, I was suggested to avoid writing one output/checkpoint file per MPI process and instead collecting data from multiple processes before outputting/checkpointing happens. I found the combination of parameters
IO::out_mode = "np" IO::out_proc_every = 8
does the job for output files, but I still have one checkpoint file per process. Is there a similar parameter, or combination of parameters, which can be used for checkpoint files?
Thank you very much, Lorenzo Ennoggi
Hello Lorenzo,
Unfortunately, Carpet will always write one checkpoint file per MPI rank, there is no way to change that.
As you learned the option out_proc_every only affects the out3D_vars output (and possible out_vars 3D output) but never checkpoints.
In my opinion, you should be impossible to stress the file system, of a reasonably provisioned cluster, with the checkpoints. Even when running on 32k MPI ranks (and 4k nodes) on BW, checkpoint-recovery was very quick (1min or so) and barely made a blip on the system monitoring radar. Any cluster with sufficiently many nodes to run at scale at 1 file per rank (for a sane number of ranks ie some OpenMP threads) should have a file system capable of taking checkpoints. Of course 1 rank per core is no longer "sane" once you go beyond a couple hundred cores.
Now writing 1 file per output variable and per MPI rank may be a different thing.... In that case out_proc_every should help with out3D_vars. I would also suggest one_file_per_group or even one_file_per_rank for this (see CarpetIOHDF5's param.ccl), which will have less of a performance (no communication) impact than out_proc_every != 1.
If the issue is opening many files (again, only for out3D_vars regular output), then you may also see benefits from the different options in:
https://bitbucket.org/eschnett/carpet/pull-requests/34
https://bitbucket.org/einsteintoolkit/tickets/issues/2364
Yours, Roland
Hello, In order to avoid stressing the filesystem on the cluster I'm running on, I was suggested to avoid writing one output/checkpoint file per MPI process and instead collecting data from multiple processes before outputting/checkpointing happens. I found the combination of parameters
IO::out_mode = "np" IO::out_proc_every = 8
does the job for output files, but I still have one checkpoint file per process. Is there a similar parameter, or combination of parameters, which can be used for checkpoint files?
Thank you very much, Lorenzo Ennoggi
Hi Roland, thank you, your suggestions are very useful. I was running one process per core on more than 200 cores, so that may be part of the issue. Also, I will try the one_file_per_group or one_file_per_rank options to reduce the performance impact.
The cluster I'm running on is Frontera, and the guidelines to manage I/O operations properly on it are here https://portal.tacc.utexas.edu/tutorials/managingio in case people are interested. I will follow them as closely as I can to avoid similar problems in the future.
Thank you very much again, Lorenzo
Il giorno gio 6 ott 2022 alle ore 12:52 Roland Haas rhaas@illinois.edu ha scritto:
Hello Lorenzo,
Unfortunately, Carpet will always write one checkpoint file per MPI rank, there is no way to change that.
As you learned the option out_proc_every only affects the out3D_vars output (and possible out_vars 3D output) but never checkpoints.
In my opinion, you should be impossible to stress the file system, of a reasonably provisioned cluster, with the checkpoints. Even when running on 32k MPI ranks (and 4k nodes) on BW, checkpoint-recovery was very quick (1min or so) and barely made a blip on the system monitoring radar. Any cluster with sufficiently many nodes to run at scale at 1 file per rank (for a sane number of ranks ie some OpenMP threads) should have a file system capable of taking checkpoints. Of course 1 rank per core is no longer "sane" once you go beyond a couple hundred cores.
Now writing 1 file per output variable and per MPI rank may be a different thing.... In that case out_proc_every should help with out3D_vars. I would also suggest one_file_per_group or even one_file_per_rank for this (see CarpetIOHDF5's param.ccl), which will have less of a performance (no communication) impact than out_proc_every != 1.
If the issue is opening many files (again, only for out3D_vars regular output), then you may also see benefits from the different options in:
https://bitbucket.org/eschnett/carpet/pull-requests/34
https://bitbucket.org/einsteintoolkit/tickets/issues/2364
Yours, Roland
Hello, In order to avoid stressing the filesystem on the cluster I'm running
on, I
was suggested to avoid writing one output/checkpoint file per MPI process and instead collecting data from multiple processes before outputting/checkpointing happens. I found the combination of parameters
IO::out_mode = "np" IO::out_proc_every = 8
does the job for output files, but I still have one checkpoint file per process. Is there a similar parameter, or combination of parameters,
which
can be used for checkpoint files?
Thank you very much, Lorenzo Ennoggi
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
Lorenzo
The thorn `Carpet/CarpetSimulationIO` can write N output files on M processes, and uses an efficient mechanism to map in between. That is, it uses the high-speed interconnect instead of disk I/O to exchange data.
This thorn is, to my knowledge, not used in production, so you'd want to test it before you use it. It needs the `SimulationIO` external library.
-erik
On Thu, Oct 6, 2022 at 12:28 PM Lorenzo Ennoggi lorenzo.ennoggi@gmail.com wrote:
Hello, In order to avoid stressing the filesystem on the cluster I'm running on, I was suggested to avoid writing one output/checkpoint file per MPI process and instead collecting data from multiple processes before outputting/checkpointing happens. I found the combination of parameters
IO::out_mode = "np" IO::out_proc_every = 8
does the job for output files, but I still have one checkpoint file per process. Is there a similar parameter, or combination of parameters, which can be used for checkpoint files?
Thank you very much, Lorenzo Ennoggi _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Hi Erik, thank you very much for getting back to me. I'll have a look a that and try to test it.
Thanks again for your help, Lorenzo
On Tue, Nov 1, 2022, 13:49 Erik Schnetter schnetter@gmail.com wrote:
Lorenzo
The thorn `Carpet/CarpetSimulationIO` can write N output files on M processes, and uses an efficient mechanism to map in between. That is, it uses the high-speed interconnect instead of disk I/O to exchange data.
This thorn is, to my knowledge, not used in production, so you'd want to test it before you use it. It needs the `SimulationIO` external library.
-erik
On Thu, Oct 6, 2022 at 12:28 PM Lorenzo Ennoggi lorenzo.ennoggi@gmail.com wrote:
Hello, In order to avoid stressing the filesystem on the cluster I'm running
on, I was suggested to avoid writing one output/checkpoint file per MPI process and instead collecting data from multiple processes before outputting/checkpointing happens. I found the combination of parameters
IO::out_mode = "np" IO::out_proc_every = 8
does the job for output files, but I still have one checkpoint file per
process. Is there a similar parameter, or combination of parameters, which can be used for checkpoint files?
Thank you very much, Lorenzo Ennoggi _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Erik Schnetter schnetter@gmail.com http://www.perimeterinstitute.ca/personal/eschnetter/
Also, I have just realized I missed the last email from Roland. I'll also have a look at branch rhaas/mpiio of Carpet and try to setup a test on few nodes to see if/how much using the parameter CarpetIOHDF5::user_MPIIO reduces the peak in the amount of data written to disk at the same time (which was the actual issue).
Thank you very much Roland, Lorenzo
On Tue, Nov 1, 2022, 23:59 Lorenzo Ennoggi lorenzo.ennoggi@gmail.com wrote:
Hi Erik, thank you very much for getting back to me. I'll have a look a that and try to test it.
Thanks again for your help, Lorenzo
On Tue, Nov 1, 2022, 13:49 Erik Schnetter schnetter@gmail.com wrote:
Lorenzo
The thorn `Carpet/CarpetSimulationIO` can write N output files on M processes, and uses an efficient mechanism to map in between. That is, it uses the high-speed interconnect instead of disk I/O to exchange data.
This thorn is, to my knowledge, not used in production, so you'd want to test it before you use it. It needs the `SimulationIO` external library.
-erik
On Thu, Oct 6, 2022 at 12:28 PM Lorenzo Ennoggi lorenzo.ennoggi@gmail.com wrote:
Hello, In order to avoid stressing the filesystem on the cluster I'm running
on, I was suggested to avoid writing one output/checkpoint file per MPI process and instead collecting data from multiple processes before outputting/checkpointing happens. I found the combination of parameters
IO::out_mode = "np" IO::out_proc_every = 8
does the job for output files, but I still have one checkpoint file per
process. Is there a similar parameter, or combination of parameters, which can be used for checkpoint files?
Thank you very much, Lorenzo Ennoggi _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Erik Schnetter schnetter@gmail.com http://www.perimeterinstitute.ca/personal/eschnetter/
users@lists.einsteintoolkit.org