Hi all. I'm having checkpoint/recovery issues with a particular simulation:
An initial short run stopped some time after iteration 32000, leaving me with checkpoints at it 30000 & 32000. I found I couldn't recover from the later of these, but as the earlier one *did* allow recovery, I didn't worry too much about it.
Now the recovered run went until some time after it 124000. I again have two sets of checkpoint data, from it 122000 and 124000. *Neither* of these work. I could imagine the later one being corrupted somehow because of disk space issues, but both?
In each case, the error output in the STDERR consists of multiple instances of the message below.
* Is this likely due to file corruption?
* What's the best way to check CarpetIOHDF5 files for corruption?
* Can I do anything about this particular run, apart from start (again) from the "good" 30000 checkpoint?
Thanks,
Bernard
---------------------------------------------------------------------------------------- HDF5-DIAG: Error detected in HDF5 (1.8.8) thread 0: #000: H5Gdeprec.c line 875 in H5Gget_objinfo(): cannot stat object major: Invalid arguments to routine minor: Unable to initialize object #001: H5Gdeprec.c line 1002 in H5G_get_objinfo(): name doesn't exist major: Symbol table minor: Object already exists #002: H5Gtraverse.c line 861 in H5G_traverse(): internal path traversal failed major: Symbol table minor: Object not found #003: H5Gtraverse.c line 641 in H5G_traverse_real(): traversal operator failed major: Symbol table minor: Callback failed #004: H5Gdeprec.c line 926 in H5G_get_objinfo_cb(): unable to get object info major: Object header minor: Can't get value #005: H5O.c line 2789 in H5O_get_info(): unable to load object header major: Object header minor: Unable to protect metadata #006: H5O.c line 1682 in H5O_protect(): unable to load object header major: Object header minor: Unable to protect metadata #007: H5AC.c line 1322 in H5AC_protect(): H5C_protect() failed. major: Object cache minor: Unable to protect metadata #008: H5C.c line 3567 in H5C_protect(): can't load entry major: Object cache minor: Unable to load metadata into cache #009: H5C.c line 7957 in H5C_load_entry(): unable to load entry major: Object cache minor: Unable to load metadata into cache #010: H5Ocache.c line 275 in H5O_load(): bad object header version number major: Object header minor: Wrong version number WARNING level 1 from host r421i1n0.p4.nas.nasa.gov process 20 while executing schedule bin CCTK_RECOVER_VARIABLES, routine IOUtil::IOUtil_RecoverGH in thorn CarpetIOHDF5, file /nobackupp8/bjkelly1/codes/Cactus.ET_2015_05/configs/vacuum/build/CarpetIOHDF5/Input.cc:1235: -> HDF5 call 'H5Gget_objinfo (group, objectname, 0, &object_info)' returned error code -1 --------------------------------------------------------------------------------------------------
Hi,
On Mon, Feb 01, 2016 at 03:39:25PM -0500, Bernard Kelly wrote:
In each case, the error output in the STDERR consists of multiple instances of the message below.
- Is this likely due to file corruption?
I would think so. A way to see would be to use another HDF5 tool to look at the file.
- What's the best way to check CarpetIOHDF5 files for corruption?
I don't know of a 'best' way, but I would first try 'h5ls' and see if already that has problems. If this succeeds, 'h5dump' might be a quick-and-dirty solution. Dump the complete file to /dev/null and see if h5dump complains about something.
- Can I do anything about this particular run, apart from start
(again) from the "good" 30000 checkpoint?
Assuming it is hdf5 file corruption, most likely not. Depending on how desperate you are you could try to see which parts are affected, and if the remaining 'good' parts are sufficient to restart your particular simulation. I wouldn't have high hopes though.
Something else that I would like to know: do you still have stdout/err of the run producing these files? Did Carpet complain during checkpoint write? If so, it might be fine continuing (as it apparently did), but shouldn't delete the last-good checkpoint file - maybe even by deleting the known-to-be-bad attempt.
Frank
Thanks, Frank.
I've looked through the combined STDOUT & STDERR file from the run that generated the 122000 & 124000 checkpoints, and it looks totally normal at that time; no complaints.
As I have N separate checkpoint files for a time step (works faster than one large one), the h5ls takes a bit of time. I'm running it now ...
This machine has been pretty flaky with its nobackup filesystem in recent weeks. I'd have been pretty unlucky to have been affected by it both at 122000 *and* 124000 (a few hours later), but it's not impossible.
Beany
On 1 February 2016 at 15:56, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi,
On Mon, Feb 01, 2016 at 03:39:25PM -0500, Bernard Kelly wrote:
In each case, the error output in the STDERR consists of multiple instances of the message below.
- Is this likely due to file corruption?
I would think so. A way to see would be to use another HDF5 tool to look at the file.
- What's the best way to check CarpetIOHDF5 files for corruption?
I don't know of a 'best' way, but I would first try 'h5ls' and see if already that has problems. If this succeeds, 'h5dump' might be a quick-and-dirty solution. Dump the complete file to /dev/null and see if h5dump complains about something.
- Can I do anything about this particular run, apart from start
(again) from the "good" 30000 checkpoint?
Assuming it is hdf5 file corruption, most likely not. Depending on how desperate you are you could try to see which parts are affected, and if the remaining 'good' parts are sufficient to restart your particular simulation. I wouldn't have high hopes though.
Something else that I would like to know: do you still have stdout/err of the run producing these files? Did Carpet complain during checkpoint write? If so, it might be fine continuing (as it apparently did), but shouldn't delete the last-good checkpoint file - maybe even by deleting the known-to-be-bad attempt.
Frank
Hello all,
This machine has been pretty flaky with its nobackup filesystem in recent weeks. I'd have been pretty unlucky to have been affected by it both at 122000 *and* 124000 (a few hours later), but it's not impossible.
By default, Cactus does not abort a run on IO errors (though HDF5 tends to scream quite loudly if there are errors). You may want to set:
IO::abort_on_io_errors = "yes"
the default handling is to print a warning.
Yours, Roland
On 1 Feb 2016, at 21:39, Bernard Kelly physicsbeany@gmail.com wrote:
Hi all. I'm having checkpoint/recovery issues with a particular simulation:
An initial short run stopped some time after iteration 32000, leaving me with checkpoints at it 30000 & 32000. I found I couldn't recover from the later of these, but as the earlier one *did* allow recovery, I didn't worry too much about it.
Now the recovered run went until some time after it 124000. I again have two sets of checkpoint data, from it 122000 and 124000. *Neither* of these work. I could imagine the later one being corrupted somehow because of disk space issues, but both?
In each case, the error output in the STDERR consists of multiple instances of the message below.
Is this likely due to file corruption?
What's the best way to check CarpetIOHDF5 files for corruption?
Hi Bernard,
There is a tool called h5check (Google for "What's the best way to check HDF5 files for corruption"):
• h5check: A tool to check the validity of an HDF5 file.
The HDF5 Format Checker, h5check, is a validation tool for verifying that an HDF5 file is encoded according to the HDF File Format Specification. Its purpose is to ensure data model integrity and long-term compatibility between evolving versions of the HDF5 library.
Note that h5check is designed and implemented without any use of the HDF5 Library.
Given a file, h5check scans through the encoded content, verifying it against the defined library format. If it finds any non-compliance, h5check prints the error and the reason behind the non-compliance; if possible, it continues the scanning. If h5check does not find any non-compliance, it prints an approval statement upon completion.
By default, the file is verified against the latest version of the file format, but the format version can be specified.
I have used this successfully in the past.
- Can I do anything about this particular run, apart from start
(again) from the "good" 30000 checkpoint?
If the file is corrupt, then I doubt it. You might be able to add debugging code to work out which dataset is corrupt, and if it is not an important one, you might be able to create a new HDF5 file with a corrected version. But this is a lot of work, and if there is more than one corrupt dataset, it's unlikely to be practical. It's probably much more realistic to just repeat the run. However, if you got corruption twice already, I suspect you will get it again. It's probably a good idea to checkpoint more frequently, as you can probably expect more corruption.
The only legitimate reason for the files being corrupted is if you ran out of disk space during write (and this is only legitimate because HDF5 does not support journaled writing, which is disappointing in 2016). If that happened, I would expect to see evidence of it in stdout/stderr, which you didn't see. If you don't have abort_on_io_errors set, then Cactus would have happily continued on after the HDF5 disk write failed, but I think it would have crashed if it couldn't write the error message to stdout/stderr, so I don't think you ran out of disk space. If the files are corrupt, it would either be a problem with the filesystem, or a bug in Cactus.
You might want to run some filesystem-checking program to see if this can be reproduced in a test case, or ask the system admins to do so.
Thanks, Ian. I got h5check (not part of the default HDF5 installation), and ran it. Each of the troublesome checkpoint does, indeed, have at least one or two "non-compliant" files there. Irritating, but I suppose it answers my question.
Roland, thanks for the abort_on_io_errors suggestion (though it might not have helped here, given the lack of warnings).
I guess I'll be starting from the older checkpoint, then.
Bernard
On 2 February 2016 at 04:30, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 1 Feb 2016, at 21:39, Bernard Kelly physicsbeany@gmail.com wrote:
Hi all. I'm having checkpoint/recovery issues with a particular simulation:
An initial short run stopped some time after iteration 32000, leaving me with checkpoints at it 30000 & 32000. I found I couldn't recover from the later of these, but as the earlier one *did* allow recovery, I didn't worry too much about it.
Now the recovered run went until some time after it 124000. I again have two sets of checkpoint data, from it 122000 and 124000. *Neither* of these work. I could imagine the later one being corrupted somehow because of disk space issues, but both?
In each case, the error output in the STDERR consists of multiple instances of the message below.
Is this likely due to file corruption?
What's the best way to check CarpetIOHDF5 files for corruption?
Hi Bernard,
There is a tool called h5check (Google for "What's the best way to check HDF5 files for corruption"):
• h5check: A tool to check the validity of an HDF5 file.
The HDF5 Format Checker, h5check, is a validation tool for verifying that an HDF5 file is encoded according to the HDF File Format Specification. Its purpose is to ensure data model integrity and long-term compatibility between evolving versions of the HDF5 library.
Note that h5check is designed and implemented without any use of the HDF5 Library.
Given a file, h5check scans through the encoded content, verifying it against the defined library format. If it finds any non-compliance, h5check prints the error and the reason behind the non-compliance; if possible, it continues the scanning. If h5check does not find any non-compliance, it prints an approval statement upon completion.
By default, the file is verified against the latest version of the file format, but the format version can be specified.
I have used this successfully in the past.
- Can I do anything about this particular run, apart from start
(again) from the "good" 30000 checkpoint?
If the file is corrupt, then I doubt it. You might be able to add debugging code to work out which dataset is corrupt, and if it is not an important one, you might be able to create a new HDF5 file with a corrected version. But this is a lot of work, and if there is more than one corrupt dataset, it's unlikely to be practical. It's probably much more realistic to just repeat the run. However, if you got corruption twice already, I suspect you will get it again. It's probably a good idea to checkpoint more frequently, as you can probably expect more corruption.
The only legitimate reason for the files being corrupted is if you ran out of disk space during write (and this is only legitimate because HDF5 does not support journaled writing, which is disappointing in 2016). If that happened, I would expect to see evidence of it in stdout/stderr, which you didn't see. If you don't have abort_on_io_errors set, then Cactus would have happily continued on after the HDF5 disk write failed, but I think it would have crashed if it couldn't write the error message to stdout/stderr, so I don't think you ran out of disk space. If the files are corrupt, it would either be a problem with the filesystem, or a bug in Cactus.
You might want to run some filesystem-checking program to see if this can be reproduced in a test case, or ask the system admins to do so.
-- Ian Hinder http://members.aei.mpg.de/ianhin
users@lists.einsteintoolkit.org