Hi all,
Thanks for fast replys ^^
Hi,
SphericalHarmonicDecomp should not be writing output at the
same time as a checkpoint. Are you using an NFS mount?
I noticed
issues with NFS servers becoming unresponsive (due to a large number
of blocking io operations) during a checkpoint. Perhaps right after
a checkpoint, the server is still too busy.
JakobI have not heard about such a problem before.When an HDF5 file is not properly closed, its content may be corrupted. (This will be addressed in the next major release.) There may be two reasons for this: either the file is not closed (which would be an error in the code), or there is a write error (e.g. you run out of disk space). The latter is the major reason for people encountering corrupted HDF5 files. Since you don't see error messages, this is either not the case, or these HDF5 output routines suppress these errors.The thorn SphericalHarmonicDecomp implements its own HDF5 output routines and does not use Cactus. I see that it uses a non-standard way to determine whether the file exists, and that it does not check for errors when writing or closing. I think that HDF5 errors should cause prominent warnings in stdout and stderr (did you check?), and if you don't see these, the writing should have succeeded.
You mention checkpointing. Are you experiencing these problems right after recovery, i.e. during the first SphericalHarmonicDecomp HDF5 output afterwards?
In this case, did you maybe switch to a new directory where this file doesn't exist?If not, then it may be the non-standard way in which the code determines whether the file already exists, combined with something that may be special about your file system.(The "standard" way operates as follows: open the file as if it existed; if this fails, open it by creating it. The code works differently: it opens the file as binary file. If this fails, the HDF5 file is created; if it succeeds, the file is closed and re-openend as HDF5 file. Maybe the quick closing-then-reopening causes problems?)-erik
Hello all,
I believe Nick Taylor at Caltech had similar issues. Bela has since
>> I have not heard about such a problem before.
fixed some bugs but had trouble actually committing them (he just saw
your emails). I'll grab his changes and commit them.