Hi all,
Here are the minutes for today’s meeting.
Cheers, Leo
------------
Chair : Sam Cupp Notes : Leo Werneck Present: Leo Werneck, Peter Diener, Sam Cupp, Steve Brandt
- Any updates on github & Ian Hindler's account? * Something was compromised with his account, but github canceled the token and looks like nothing bad happened. Ian is working on it.
- ET mailing list archives not using good SSL certificates * Steve will contact Roland, since this was not discussed last week but was in the agenda.
- Safety feature to avoid HDF5 files from being corrupted * Leo requests a feature that would allow the user to e.g., generate one output file per restart. With kuibit, there was interested in switching from the ASCII data files to the HDF5 in our research group. However, in a recent simulation it turned out that a node failure caused a crash as one of the HDF5 was being written to and we lost all data for an important gridfunction. If one HDF5 file was written per restart (or another safety feature was in place), then this would have not been an issue, as only one of the chunks of data would have been corrupted. Leo will open a ticket about this.
- Unanswered questions on the mailing list * Spandan Sarma had an issue where it seemed the code didn't scale as expected, but Peter pointed out that the scaling is more or less what's expected.
- Open tickets * All website related tickets are likely more interesting to Roland, so we won't be discussing them today.
Next meeting : possibly Jan 5, 2023 Next chair : pending Next minute taker: pending
------ Leonardo R. Werneck, Ph.D. Postdoctoral researcher Office EP 314 | Department of Physics | University of Idaho 875 Perimeter Dr. MS 0903 Moscow, ID 83844-0903, USA leonardo@uidaho.edu mailto:leonardo@uidaho.edu https://leowerneck.github.io https://leowerneck.github.io/
- Safety feature to avoid HDF5 files from being corrupted
- Leo requests a feature that would allow the user to e.g., generate one output file per restart. With kuibit, there was interested in
switching from the ASCII data files to the HDF5 in our research group. However, in a recent simulation it turned out that a node failure caused a crash as one of the HDF5 was being written to and we lost all data for an important gridfunction. If one HDF5 file was written per restart (or another safety feature was in place), then this would have not been an issue, as only one of the chunks of data would have been corrupted. Leo will open a ticket about this.
Isn't this done automatically when using simfactory? I have my hdf5 data written in the separate output-00?? directories (the ones generated by symfactory at each restart) so that if one run has problems I do not lose all the data.
Cheers, Bruno
I don't think this is by default per say. I use batchtools (instead of simfactory) exactly for this reason and discourage new users from restarting in the same directory as the parent checkpoints to avoid this exact outcome. An additional issue that can arise is if a job is terminated before the walltime such that the data stored in ASCII/HDF5 goes beyond the last checkpoint are potential ingestion issues due to data mismatching from the restart. I think overall Kuibit handles this well, but it is has been an issue in the past for some users before learning to separate restarts into individual directories.
Cheers,
Samuel
On 12/16/22 10:01 AM, Bruno Giacomazzo wrote:
- Safety feature to avoid HDF5 files from being corrupted * Leo requests a feature that would allow the user to e.g., generate one output file per restart. With kuibit, there was interested in switching from the ASCII data files to the HDF5 in our research group. However, in a recent simulation it turned out that a node failure caused a crash as one of the HDF5 was being written to and we lost all data for an important gridfunction. If one HDF5 file was written per restart (or another safety feature was in place), then this would have not been an issue, as only one of the chunks of data would have been corrupted. Leo will open a ticket about this.Isn't this done automatically when using simfactory? I have my hdf5 data written in the separate output-00?? directories (the ones generated by symfactory at each restart) so that if one run has problems I do not lose all the data.
Cheers, Bruno
--
Prof. Bruno Giacomazzo Department of Physics University of Milano-Bicocca Piazza della Scienza 3 20126 Milano Italy
email: bruno.giacomazzo@unimib.it phone: (+39) 02 6448 2321 web: http://www.brunogiacomazzo.org
There are only 10 types of people in the world: Those who understand binary, and those who don't
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
When using simfactory, if I write in my parameter file that the output directory for the hdf5 files is "./hdf5" (e.g., CarpetIOHDF5::out2D_dir = "./hdf5_2D") then this directory is created in the output-00?? directory. Therefore each run creates a separate hdf5 directory in the corresponding output-00?? one. This is how I avoid losing all data if one run has problems.
Cheers, Bruno
Il giorno ven 16 dic 2022 alle ore 10:10 Samuel Tootle < tootle@itp.uni-frankfurt.de> ha scritto:
I don't think this is by default per say. I use batchtools (instead of simfactory) exactly for this reason and discourage new users from restarting in the same directory as the parent checkpoints to avoid this exact outcome. An additional issue that can arise is if a job is terminated before the walltime such that the data stored in ASCII/HDF5 goes beyond the last checkpoint are potential ingestion issues due to data mismatching from the restart. I think overall Kuibit handles this well, but it is has been an issue in the past for some users before learning to separate restarts into individual directories.
Cheers,
Samuel On 12/16/22 10:01 AM, Bruno Giacomazzo wrote:
- Safety feature to avoid HDF5 files from being corrupted
- Leo requests a feature that would allow the user to e.g., generate one output file per restart. With kuibit, there was interested in
switching from the ASCII data files to the HDF5 in our research group. However, in a recent simulation it turned out that a node failure caused a crash as one of the HDF5 was being written to and we lost all data for an important gridfunction. If one HDF5 file was written per restart (or another safety feature was in place), then this would have not been an issue, as only one of the chunks of data would have been corrupted. Leo will open a ticket about this.
Isn't this done automatically when using simfactory? I have my hdf5 data written in the separate output-00?? directories (the ones generated by symfactory at each restart) so that if one run has problems I do not lose all the data.
Cheers, Bruno
--
Prof. Bruno Giacomazzo Department of Physics University of Milano-Bicocca Piazza della Scienza 3 20126 Milano Italy
email: bruno.giacomazzo@unimib.it phone: (+39) 02 6448 2321 web: http://www.brunogiacomazzo.org
There are only 10 types of people in the world: Those who understand binary, and those who don't
Users mailing listUsers@einsteintoolkit.orghttp://lists.einsteintoolkit.org/mailman/listinfo/users
Bruno
Yes, this was one of the main design goals of the Simulation Factory. Many things can go wrong on supercomputers, and Simfactory automates what one should do to be safe and efficient.
Of course, people can circumvent this safety feature by escaping upwards with the output directory or by using hard-coded paths. Please don't do this. Set the output directory to "$parfile".
-erik
On Fri, Dec 16, 2022 at 4:01 AM Bruno Giacomazzo bruno.giacomazzo@unimib.it wrote:
- Safety feature to avoid HDF5 files from being corrupted
- Leo requests a feature that would allow the user to e.g., generate one output file per restart. With kuibit, there was interested in switching from the ASCII data files to the HDF5 in our research group. However, in a recent simulation it turned out that a node failure caused a crash as one of the HDF5 was being written to and we lost all data for an important gridfunction. If one HDF5 file was written per restart (or another safety feature was in place), then this would have not been an issue, as only one of the chunks of data would have been corrupted. Leo will open a ticket about this.
Isn't this done automatically when using simfactory? I have my hdf5 data written in the separate output-00?? directories (the ones generated by symfactory at each restart) so that if one run has problems I do not lose all the data.
Cheers, Bruno
--
Prof. Bruno Giacomazzo Department of Physics University of Milano-Bicocca Piazza della Scienza 3 20126 Milano Italy
email: bruno.giacomazzo@unimib.it phone: (+39) 02 6448 2321 web: http://www.brunogiacomazzo.org
There are only 10 types of people in the world: Those who understand binary, and those who don't
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
users@lists.einsteintoolkit.org