Hi,
I am having problems with checkpoint restarting for high resolution Carper/McLachlan simulations. For more coarse resolutions I have no problems with checkpoint restarting, nor do I have any problems if the high resolution simulations run for only a few M. But when I do high-resolution simulations until stopped by wall-time and try to restart from checkpoint, I get the following error output :
################################################################# HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5Dio.c line 174 in H5Dread(): can't read data major: Dataset minor: Read failed #001: H5Dio.c line 404 in H5D_read(): can't read data major: Dataset minor: Read failed #002: H5Dchunk.c line 1724 in H5D_chunk_read(): unable to read raw data chunk major: Low-level I/O minor: Read failed #003: H5Dchunk.c line 2737 in H5D_chunk_lock(): data pipeline read failed major: Data filters minor: Filter operation failed #004: H5Z.c line 1116 in H5Z_pipeline(): filter returned failure during read major: Data filters minor: Read failed #005: H5Zdeflate.c line 133 in H5Z_filter_deflate(): memory allocation failed for deflate uncompression major: Resource unavailable minor: No space available for allocation WARNING level 1 in thorn CarpetIOHDF5 processor 104 host tachyon3167 (line 1102 of /home01/r632kgw/jakob/Cactus/arrangements/Carpet/CarpetIOHDF5/src/Input.cc):
-> HDF5 call 'H5Dread (dataset, datatype, memspace, filespace, xfer, cctkGH->data[patch->vindex][timelevel])' returned error code -1 . . ......etc, etc, ##################################################################
Any ideas for possible cause and solution to this?
Thanks, Jakob
Hi,
On Thu, Feb 17, 2011 at 03:24:04PM +0900, Jakob Hansen wrote:
#005: H5Zdeflate.c line 133 in H5Z_filter_deflate(): memory allocation failed for deflate uncompression major: Resource unavailable minor: No space available for allocation
Any ideas for possible cause and solution to this?
Looking at these error messages I suggest you first do an h5ls/h5dump to see if the files are actually ok. If that is so, what I would suspect next is that reading the files takes more memory than is available, as indicated by the message above. In this case I suggest to try the workarounds mentioned in this thread:
http://lists.einsteintoolkit.org/pipermail/users/2011-February/000852.html
If this also doesn't help I would suggest to try again using more available memory, e.g. using more nodes.
Frank
Hi Jakob,
I had the same problem some time ago. Since your HDF5 files are compressed, Cactus tries to unpack them before reading, and this operation somehow is very memory-intensive. Try to repack your checkpoint files before restarting to reduce compression level to zero. You can use the standard h5repack utility for that. Here's a small script to uncompress all checkpoint files and put them to an $OUTDIR directory:
---------------- #!/bin/bash
ITER=5555 # which iteration to use OUTDIR=$SCRATCH/tmp # output directory for f in checkpoint.chkpt.it_${ITER}.file_*.h5; do echo processing $f... h5repack -i $f -o $OUTDIR/$f -f NONE done ----------------
This will make your checkpoint files 5-10 times larger, but because they don't need to be unpacked in RAM, less memory is required to read them.
Also, try the following: 1. set: CarpetIOHDF5::open_one_input_file_at_a_time = yes IO::abort_on_io_errors = yes 2. use less cores per node. 3. try running on larger number of processors.
Cheers, - Oleg Korobkin
Frank Loeffler wrote:
Hi,
On Thu, Feb 17, 2011 at 03:24:04PM +0900, Jakob Hansen wrote:
#005: H5Zdeflate.c line 133 in H5Z_filter_deflate(): memory allocation failed for deflate uncompression major: Resource unavailable minor: No space available for allocation
Any ideas for possible cause and solution to this?
Looking at these error messages I suggest you first do an h5ls/h5dump to see if the files are actually ok. If that is so, what I would suspect next is that reading the files takes more memory than is available, as indicated by the message above. In this case I suggest to try the workarounds mentioned in this thread:
http://lists.einsteintoolkit.org/pipermail/users/2011-February/000852.html
If this also doesn't help I would suggest to try again using more available memory, e.g. using more nodes.
Frank
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Thank you for suggestions,
Oleg, your h5repack script did the trick and I can now resume my simulation, thanks :)
Cheers, Jakob
2011/2/18 Oleg Korobkin korobkin@phys.lsu.edu
Hi Jakob,
I had the same problem some time ago. Since your HDF5 files are compressed, Cactus tries to unpack them before reading, and this operation somehow is very memory-intensive. Try to repack your checkpoint files before restarting to reduce compression level to zero. You can use the standard h5repack utility for that. Here's a small script to uncompress all checkpoint files and put them to an $OUTDIR directory:
#!/bin/bash
ITER=5555 # which iteration to use OUTDIR=$SCRATCH/tmp # output directory for f in checkpoint.chkpt.it_${ITER}.file_*.h5; do echo processing $f... h5repack -i $f -o $OUTDIR/$f -f NONE done
This will make your checkpoint files 5-10 times larger, but because they don't need to be unpacked in RAM, less memory is required to read them.
Also, try the following:
- set:
CarpetIOHDF5::open_one_input_file_at_a_time = yes IO::abort_on_io_errors = yes 2. use less cores per node. 3. try running on larger number of processors.
Cheers,
- Oleg Korobkin
Frank Loeffler wrote:
Hi,
On Thu, Feb 17, 2011 at 03:24:04PM +0900, Jakob Hansen wrote:
#005: H5Zdeflate.c line 133 in H5Z_filter_deflate(): memory allocation failed for deflate uncompression major: Resource unavailable minor: No space available for allocation
Any ideas for possible cause and solution to this?
Looking at these error messages I suggest you first do an h5ls/h5dump to see if the files are actually ok. If that is so, what I would suspect
next
is that reading the files takes more memory than is available, as
indicated
by the message above. In this case I suggest to try the workarounds mentioned in this thread:
http://lists.einsteintoolkit.org/pipermail/users/2011-February/000852.html
If this also doesn't help I would suggest to try again using more available memory, e.g. using more nodes.
Frank
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
users@lists.einsteintoolkit.org