Hi,
I'm experimenting a bit with using the PittNull / SphericalHarmonicDecomp thorn but I've been experiencing a few errors in writing to its output metric_obs_0_Decomp.h5 file.
On some occations it seems that this output file cannot be opened and hence leaving gaps in the data file. This always occurs when carpet is doing checkpoininting. Everything else runs well, though, and the simulations finishes fine. However, when trying to fftwfilter the metric_obs_0_Decomp.h5 file, I notice the missing data points.
I've checked system logs and there seem to be no hardware failure. Also, as far as I can see, I'm not overusing memory allocation either.
Anyone else experienced similar issues?
Here follows output from error file : -----------------------------------------------------------------------------------------
HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5F.c line 1509 in H5Fopen(): unable to open file major: File accessability minor: Unable to open file #001: H5F.c line 1300 in H5F_open(): unable to read superblock major: File accessability minor: Read failed #002: H5Fsuper.c line 324 in H5F_super_read(): unable to load superblock major: Object cache minor: Unable to protect metadata #003: H5AC.c line 1597 in H5AC_protect(): H5C_protect() failed. major: Object cache minor: Unable to protect metadata #004: H5C.c line 3333 in H5C_protect(): can't load entry major: Object cache minor: Unable to load metadata into cache #005: H5C.c line 8177 in H5C_load_entry(): unable to load entry major: Object cache minor: Unable to load metadata into cache #006: H5Fsuper_cache.c line 469 in H5F_sblock_load(): truncated file major: File accessability minor: File has been truncated HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5Gdeprec.c line 214 in H5Gcreate1(): not a location major: Invalid arguments to routine minor: Inappropriate type #001: H5Gloc.c line 253 in H5G_loc(): invalid object ID major: Invalid arguments to routine minor: Bad value HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5Adeprec.c line 153 in H5Acreate1(): not a location major: Invalid arguments to routine minor: Inappropriate type #001: H5Gloc.c line 253 in H5G_loc(): invalid object ID major: Invalid arguments to routine minor: Bad value [etc. etc. etc. ...... ] -------------------------------------------
Cheers, Jakob
Hi,
SphericalHarmonicDecomp should not be writing output at the same time as a checkpoint. Are you using an NFS mount? I noticed issues with NFS servers becoming unresponsive (due to a large number of blocking io operations) during a checkpoint. Perhaps right after a checkpoint, the server is still too busy.
This certainly sounds like an issue that needs to be fixed, although I am not sure how. Perhaps a failed IO operation should be a fatal error that kills the run.
Could you try setting the output for metric_obs_0_Decomp.h5 so that it doesn't correspond to a the iteration immediately before or after a checkpoint?
On 07/27/2012 03:59 AM, Jakob Hansen wrote:
Hi,
I'm experimenting a bit with using the PittNull / SphericalHarmonicDecomp thorn but I've been experiencing a few errors in writing to its output metric_obs_0_Decomp.h5 file.
On some occations it seems that this output file cannot be opened and hence leaving gaps in the data file. This always occurs when carpet is doing checkpoininting. Everything else runs well, though, and the simulations finishes fine. However, when trying to fftwfilter the metric_obs_0_Decomp.h5 file, I notice the missing data points.
I've checked system logs and there seem to be no hardware failure. Also, as far as I can see, I'm not overusing memory allocation either.
Anyone else experienced similar issues?
Here follows output from error file :
HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5F.c line 1509 in H5Fopen(): unable to open file major: File accessability minor: Unable to open file #001: H5F.c line 1300 in H5F_open(): unable to read superblock major: File accessability minor: Read failed #002: H5Fsuper.c line 324 in H5F_super_read(): unable to load superblock major: Object cache minor: Unable to protect metadata #003: H5AC.c line 1597 in H5AC_protect(): H5C_protect() failed. major: Object cache minor: Unable to protect metadata #004: H5C.c line 3333 in H5C_protect(): can't load entry major: Object cache minor: Unable to load metadata into cache #005: H5C.c line 8177 in H5C_load_entry(): unable to load entry major: Object cache minor: Unable to load metadata into cache #006: H5Fsuper_cache.c line 469 in H5F_sblock_load(): truncated file major: File accessability minor: File has been truncated HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5Gdeprec.c line 214 in H5Gcreate1(): not a location major: Invalid arguments to routine minor: Inappropriate type #001: H5Gloc.c line 253 in H5G_loc(): invalid object ID major: Invalid arguments to routine minor: Bad value HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5Adeprec.c line 153 in H5Acreate1(): not a location major: Invalid arguments to routine minor: Inappropriate type #001: H5Gloc.c line 253 in H5G_loc(): invalid object ID major: Invalid arguments to routine minor: Bad value [etc. etc. etc. ...... ]
Cheers, Jakob
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Jakob
I have not heard about such a problem before.
When an HDF5 file is not properly closed, its content may be corrupted. (This will be addressed in the next major release.) There may be two reasons for this: either the file is not closed (which would be an error in the code), or there is a write error (e.g. you run out of disk space). The latter is the major reason for people encountering corrupted HDF5 files. Since you don't see error messages, this is either not the case, or these HDF5 output routines suppress these errors.
The thorn SphericalHarmonicDecomp implements its own HDF5 output routines and does not use Cactus. I see that it uses a non-standard way to determine whether the file exists, and that it does not check for errors when writing or closing. I think that HDF5 errors should cause prominent warnings in stdout and stderr (did you check?), and if you don't see these, the writing should have succeeded.
You mention checkpointing. Are you experiencing these problems right after recovery, i.e. during the first SphericalHarmonicDecomp HDF5 output afterwards? In this case, did you maybe switch to a new directory where this file doesn't exist?
If not, then it may be the non-standard way in which the code determines whether the file already exists, combined with something that may be special about your file system.
(The "standard" way operates as follows: open the file as if it existed; if this fails, open it by creating it. The code works differently: it opens the file as binary file. If this fails, the HDF5 file is created; if it succeeds, the file is closed and re-openend as HDF5 file. Maybe the quick closing-then-reopening causes problems?)
-erik
On Friday, July 27, 2012, Jakob Hansen wrote:
Hi,
I'm experimenting a bit with using the PittNull / SphericalHarmonicDecomp thorn but I've been experiencing a few errors in writing to its output metric_obs_0_Decomp.h5 file.
On some occations it seems that this output file cannot be opened and hence leaving gaps in the data file. This always occurs when carpet is doing checkpoininting. Everything else runs well, though, and the simulations finishes fine. However, when trying to fftwfilter the metric_obs_0_Decomp.h5 file, I notice the missing data points.
I've checked system logs and there seem to be no hardware failure. Also, as far as I can see, I'm not overusing memory allocation either.
Anyone else experienced similar issues?
Here follows output from error file :
HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5F.c line 1509 in H5Fopen(): unable to open file major: File accessability minor: Unable to open file #001: H5F.c line 1300 in H5F_open(): unable to read superblock major: File accessability minor: Read failed #002: H5Fsuper.c line 324 in H5F_super_read(): unable to load superblock major: Object cache minor: Unable to protect metadata #003: H5AC.c line 1597 in H5AC_protect(): H5C_protect() failed. major: Object cache minor: Unable to protect metadata #004: H5C.c line 3333 in H5C_protect(): can't load entry major: Object cache minor: Unable to load metadata into cache #005: H5C.c line 8177 in H5C_load_entry(): unable to load entry major: Object cache minor: Unable to load metadata into cache #006: H5Fsuper_cache.c line 469 in H5F_sblock_load(): truncated file major: File accessability minor: File has been truncated HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5Gdeprec.c line 214 in H5Gcreate1(): not a location major: Invalid arguments to routine minor: Inappropriate type #001: H5Gloc.c line 253 in H5G_loc(): invalid object ID major: Invalid arguments to routine minor: Bad value HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5Adeprec.c line 153 in H5Acreate1(): not a location major: Invalid arguments to routine minor: Inappropriate type #001: H5Gloc.c line 253 in H5G_loc(): invalid object ID major: Invalid arguments to routine minor: Bad value [etc. etc. etc. ...... ]
Cheers, Jakob
On 07/27/2012 12:45 PM, Erik Schnetter wrote:
Jakob
I have not heard about such a problem before.
When an HDF5 file is not properly closed, its content may be corrupted. (This will be addressed in the next major release.) There may be two reasons for this: either the file is not closed (which would be an error in the code), or there is a write error (e.g. you run out of disk space). The latter is the major reason for people encountering corrupted HDF5 files. Since you don't see error messages, this is either not the case, or these HDF5 output routines suppress these errors.
The thorn SphericalHarmonicDecomp implements its own HDF5 output routines and does not use Cactus. I see that it uses a non-standard way to determine whether the file exists, and that it does not check for errors when writing or closing. I think that HDF5 errors should cause prominent warnings in stdout and stderr (did you check?), and if you don't see these, the writing should have succeeded.
You mention checkpointing. Are you experiencing these problems right after recovery, i.e. during the first SphericalHarmonicDecomp HDF5 output afterwards? In this case, did you maybe switch to a new directory where this file doesn't exist?
When I ran CCE, I always restarted in a new directory and recombined the hdf5 files after the run finished.
If not, then it may be the non-standard way in which the code determines whether the file already exists, combined with something that may be special about your file system.
(The "standard" way operates as follows: open the file as if it existed; if this fails, open it by creating it. The code works differently: it opens the file as binary file. If this fails, the HDF5 file is created; if it succeeds, the file is closed and re-openend as HDF5 file. Maybe the quick closing-then-reopening causes problems?)
I probably should fix this. BTW, carpetIOHDF5 doesn't do this. Instead, it uses H5Fis_hdf5(filename) >0;. I'll try to work on cleaning the code.
-erik
On Friday, July 27, 2012, Jakob Hansen wrote:
Hi, I'm experimenting a bit with using the PittNull / SphericalHarmonicDecomp thorn but I've been experiencing a few errors in writing to its output metric_obs_0_Decomp.h5 file. On some occations it seems that this output file cannot be opened and hence leaving gaps in the data file. This always occurs when carpet is doing checkpoininting. Everything else runs well, though, and the simulations finishes fine. However, when trying to fftwfilter the metric_obs_0_Decomp.h5 file, I notice the missing data points. I've checked system logs and there seem to be no hardware failure. Also, as far as I can see, I'm not overusing memory allocation either. Anyone else experienced similar issues? Here follows output from error file : ----------------------------------------------------------------------------------------- HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5F.c line 1509 in H5Fopen(): unable to open file major: File accessability minor: Unable to open file #001: H5F.c line 1300 in H5F_open(): unable to read superblock major: File accessability minor: Read failed #002: H5Fsuper.c line 324 in H5F_super_read(): unable to load superblock major: Object cache minor: Unable to protect metadata #003: H5AC.c line 1597 in H5AC_protect(): H5C_protect() failed. major: Object cache minor: Unable to protect metadata #004: H5C.c line 3333 in H5C_protect(): can't load entry major: Object cache minor: Unable to load metadata into cache #005: H5C.c line 8177 in H5C_load_entry(): unable to load entry major: Object cache minor: Unable to load metadata into cache #006: H5Fsuper_cache.c line 469 in H5F_sblock_load(): truncated file major: File accessability minor: File has been truncated HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5Gdeprec.c line 214 in H5Gcreate1(): not a location major: Invalid arguments to routine minor: Inappropriate type #001: H5Gloc.c line 253 in H5G_loc(): invalid object ID major: Invalid arguments to routine minor: Bad value HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5Adeprec.c line 153 in H5Acreate1(): not a location major: Invalid arguments to routine minor: Inappropriate type #001: H5Gloc.c line 253 in H5G_loc(): invalid object ID major: Invalid arguments to routine minor: Bad value [etc. etc. etc. ...... ] ------------------------------------------- Cheers, Jakob-- Erik Schnetter <schnetter@cct.lsu.edu mailto:schnetter@cct.lsu.edu> http://www.perimeterinstitute.ca/personal/eschnetter/
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Hello all,
I have not heard about such a problem before.
I believe Nick Taylor at Caltech had similar issues. Bela has since fixed some bugs but had trouble actually committing them (he just saw your emails). I'll grab his changes and commit them.
Yours, Roland
Hi all,
Thanks for fast replys ^^
2012/7/27 Yosef Zlochower yosef@astro.rit.edu
Hi,
SphericalHarmonicDecomp should not be writing output at the same time as a checkpoint. Are you using an NFS mount?
No, we're using Lustre filesystem.
I noticed
issues with NFS servers becoming unresponsive (due to a large number of blocking io operations) during a checkpoint. Perhaps right after a checkpoint, the server is still too busy.
Well, indeed this happens just after checkpointing, however not at every checkpoint and not for all output. I experienced this error twice on two different simulations, once for each simulation. In each case it affected the metric_obs_0_Decomp.h5 right after checkpointing :
Case 1 : Output from ascii_output > gxx.asc : 2.7627600000000001e+02 -7.1882256506489145e-05 1.5746413856874709e-05 2.7640800000000002e+02 -7.1881310717781865e-05 1.5759232641588166e-05 nan 0.0000000000000000e+00 0.0000000000000000e+00 2.7667200000000003e+02 -7.1875929421365471e-05 1.5784711415784759e-05 2.7680400000000003e+02 -7.1871563990008602e-05 1.5797542778149928e-05
In this case there was a checkpoint at time 276.408 : INFO (CarpetIOHDF5): Dumping periodic checkpoint at iteration 268032, simulation time 276.408
Case 2: Output from ascii_output > gxx.asc : 3.2010000000000002e+02 -6.7572912816444132e-05 2.1144268516557760e-05 3.2023200000000003e+02 -6.7570803118733184e-05 2.1156692762978387e-05 nan 0.0000000000000000e+00 0.0000000000000000e+00 3.2049600000000004e+02 -6.7568973330671772e-05 2.1182079214383173e-05 3.2062800000000004e+02 -6.7569118827631724e-05 2.1194791416536239e-05
With a checkpoint at time 320.232 : INFO (CarpetIOHDF5): Dumping periodic checkpoint at iteration 310528, simulation time 320.232
Also, in both cases it only affected the metric_obs_0_Decomp.h5 file, the other detection radius files, _1 and _2, had all data.
<snip>
2012/7/28 Erik Schnetter schnetter@cct.lsu.edu
Jakob
I have not heard about such a problem before.
When an HDF5 file is not properly closed, its content may be corrupted. (This will be addressed in the next major release.) There may be two reasons for this: either the file is not closed (which would be an error in the code), or there is a write error (e.g. you run out of disk space). The latter is the major reason for people encountering corrupted HDF5 files. Since you don't see error messages, this is either not the case, or these HDF5 output routines suppress these errors.
The thorn SphericalHarmonicDecomp implements its own HDF5 output routines and does not use Cactus. I see that it uses a non-standard way to determine whether the file exists, and that it does not check for errors when writing or closing. I think that HDF5 errors should cause prominent warnings in stdout and stderr (did you check?), and if you don't see these, the writing should have succeeded.
The errors I see on stderr are the ones I mentioned in my first mail :
HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5F.c line 1509 in H5Fopen(): unable to open file major: File accessability minor: Unable to open file #001: H5F.c line 1300 in H5F_open(): unable to read superblock major: File accessability minor: Read failed
.... etc.
stdout does not produce any errors or warnings related to this.
You mention checkpointing. Are you experiencing these problems right after
recovery, i.e. during the first SphericalHarmonicDecomp HDF5 output afterwards?
No, this happened during the first run, not related to recovery.
In this case, did you maybe switch to a new directory where this file doesn't exist?
If not, then it may be the non-standard way in which the code determines whether the file already exists, combined with something that may be special about your file system.
(The "standard" way operates as follows: open the file as if it existed; if this fails, open it by creating it. The code works differently: it opens the file as binary file. If this fails, the HDF5 file is created; if it succeeds, the file is closed and re-openend as HDF5 file. Maybe the quick closing-then-reopening causes problems?)
-erik
2012/7/28 Roland Haas roland.haas@physics.gatech.edu
Hello all,
I have not heard about such a problem before.
I believe Nick Taylor at Caltech had similar issues. Bela has since fixed some bugs but had trouble actually committing them (he just saw your emails). I'll grab his changes and commit them.
Sounds interesting, I'm looking forward to appying the changes and see if the problem disappears.
Cheers, Jakob
On 07/28/2012 07:38 AM, Jakob Hansen wrote:
Hi all,
Thanks for fast replys ^^
2012/7/27 Yosef Zlochower <yosef@astro.rit.edu mailto:yosef@astro.rit.edu>
Hi, SphericalHarmonicDecomp should not be writing output at the same time as a checkpoint. Are you using an NFS mount?No, we're using Lustre filesystem.
I noticed issues with NFS servers becoming unresponsive (due to a large number of blocking io operations) during a checkpoint. Perhaps right after a checkpoint, the server is still too busy.Well, indeed this happens just after checkpointing, however not at every checkpoint and not for all output. I experienced this error twice on two different simulations, once for each simulation. In each case it affected the metric_obs_0_Decomp.h5 right after checkpointing :
Case 1 : Output from ascii_output > gxx.asc : 2.7627600000000001e+02 -7.1882256506489145e-05 1.5746413856874709e-05 2.7640800000000002e+02 -7.1881310717781865e-05 1.5759232641588166e-05 nan 0.0000000000000000e+00 0.0000000000000000e+00 2.7667200000000003e+02 -7.1875929421365471e-05 1.5784711415784759e-05 2.7680400000000003e+02 -7.1871563990008602e-05 1.5797542778149928e-05
In this case there was a checkpoint at time 276.408 : INFO (CarpetIOHDF5): Dumping periodic checkpoint at iteration 268032, simulation time 276.408
The NaN is just the way ascii_ouput let's you know it couldn't read the data.
Case 2: Output from ascii_output > gxx.asc : 3.2010000000000002e+02 -6.7572912816444132e-05 2.1144268516557760e-05 3.2023200000000003e+02 -6.7570803118733184e-05 2.1156692762978387e-05 nan 0.0000000000000000e+00 0.0000000000000000e+00 3.2049600000000004e+02 -6.7568973330671772e-05 2.1182079214383173e-05 3.2062800000000004e+02 -6.7569118827631724e-05 2.1194791416536239e-05
With a checkpoint at time 320.232 : INFO (CarpetIOHDF5): Dumping periodic checkpoint at iteration 310528, simulation time 320.232
Also, in both cases it only affected the metric_obs_0_Decomp.h5 file, the other detection radius files, _1 and _2, had all data.
Do the Caltech fixes help? If that doesn't work, then it may be that your IO system is saturated. Perhaps then a crude workaround would be to put a delay in after a checkpoint to give the IO system time to process its backlog of IO requests.
<snip>
2012/7/28 Erik Schnetter <schnetter@cct.lsu.edu mailto:schnetter@cct.lsu.edu>
Jakob I have not heard about such a problem before. When an HDF5 file is not properly closed, its content may be corrupted. (This will be addressed in the next major release.) There may be two reasons for this: either the file is not closed (which would be an error in the code), or there is a write error (e.g. you run out of disk space). The latter is the major reason for people encountering corrupted HDF5 files. Since you don't see error messages, this is either not the case, or these HDF5 output routines suppress these errors. The thorn SphericalHarmonicDecomp implements its own HDF5 output routines and does not use Cactus. I see that it uses a non-standard way to determine whether the file exists, and that it does not check for errors when writing or closing. I think that HDF5 errors should cause prominent warnings in stdout and stderr (did you check?), and if you don't see these, the writing should have succeeded.The errors I see on stderr are the ones I mentioned in my first mail :
HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5F.c line 1509 in H5Fopen(): unable to open file major: File accessability minor: Unable to open file #001: H5F.c line 1300 in H5F_open(): unable to read superblock major: File accessability minor: Read failed
.... etc.
stdout does not produce any errors or warnings related to this.
You mention checkpointing. Are you experiencing these problems right after recovery, i.e. during the first SphericalHarmonicDecomp HDF5 output afterwards?No, this happened during the first run, not related to recovery.
In this case, did you maybe switch to a new directory where this file doesn't exist? If not, then it may be the non-standard way in which the code determines whether the file already exists, combined with something that may be special about your file system. (The "standard" way operates as follows: open the file as if it existed; if this fails, open it by creating it. The code works differently: it opens the file as binary file. If this fails, the HDF5 file is created; if it succeeds, the file is closed and re-openend as HDF5 file. Maybe the quick closing-then-reopening causes problems?) -erik2012/7/28 Roland Haas <roland.haas@physics.gatech.edu mailto:roland.haas@physics.gatech.edu>
Hello all, >> I have not heard about such a problem before. I believe Nick Taylor at Caltech had similar issues. Bela has since fixed some bugs but had trouble actually committing them (he just saw your emails). I'll grab his changes and commit them.Sounds interesting, I'm looking forward to appying the changes and see if the problem disappears.
Cheers, Jakob
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
I wonder if this hack may help.
In SphericalHarmonicDecomp_DumpMetric add a blocking IO operation before the Decompose3D calls. Perhaps something like:
{ const char *outdir = *out_dir ? out_dir : io_out_dir; char filename[BUFFSIZE]; snprintf(filename, sizeof filename, "%s/obs_%d_test_io_ready", dir, obs); FILE *file = fopen(filename, "a"); assert (file); fprintf(file, "test_if_ready\n"); fflush(file); fclose(file) }
If that doesn't help, then perhaps you can set SphericalHamonicDecomp to abort the run when this happens.
On 07/29/2012 04:31 PM, Yosef Zlochower wrote:
On 07/28/2012 07:38 AM, Jakob Hansen wrote:
Hi all,
Thanks for fast replys ^^
2012/7/27 Yosef Zlochower <yosef@astro.rit.edu mailto:yosef@astro.rit.edu>
Hi, SphericalHarmonicDecomp should not be writing output at the same time as a checkpoint. Are you using an NFS mount?No, we're using Lustre filesystem.
I noticed issues with NFS servers becoming unresponsive (due to a large number of blocking io operations) during a checkpoint. Perhaps right after a checkpoint, the server is still too busy.Well, indeed this happens just after checkpointing, however not at every checkpoint and not for all output. I experienced this error twice on two different simulations, once for each simulation. In each case it affected the metric_obs_0_Decomp.h5 right after checkpointing :
Case 1 : Output from ascii_output > gxx.asc : 2.7627600000000001e+02 -7.1882256506489145e-05 1.5746413856874709e-05 2.7640800000000002e+02 -7.1881310717781865e-05 1.5759232641588166e-05 nan 0.0000000000000000e+00 0.0000000000000000e+00 2.7667200000000003e+02 -7.1875929421365471e-05 1.5784711415784759e-05 2.7680400000000003e+02 -7.1871563990008602e-05 1.5797542778149928e-05
In this case there was a checkpoint at time 276.408 : INFO (CarpetIOHDF5): Dumping periodic checkpoint at iteration 268032, simulation time 276.408
The NaN is just the way ascii_ouput let's you know it couldn't read the data.
Case 2: Output from ascii_output > gxx.asc : 3.2010000000000002e+02 -6.7572912816444132e-05 2.1144268516557760e-05 3.2023200000000003e+02 -6.7570803118733184e-05 2.1156692762978387e-05 nan 0.0000000000000000e+00 0.0000000000000000e+00 3.2049600000000004e+02 -6.7568973330671772e-05 2.1182079214383173e-05 3.2062800000000004e+02 -6.7569118827631724e-05 2.1194791416536239e-05
With a checkpoint at time 320.232 : INFO (CarpetIOHDF5): Dumping periodic checkpoint at iteration 310528, simulation time 320.232
Also, in both cases it only affected the metric_obs_0_Decomp.h5 file, the other detection radius files, _1 and _2, had all data.
Do the Caltech fixes help? If that doesn't work, then it may be that your IO system is saturated. Perhaps then a crude workaround would be to put a delay in after a checkpoint to give the IO system time to process its backlog of IO requests.
<snip>
2012/7/28 Erik Schnetter <schnetter@cct.lsu.edu mailto:schnetter@cct.lsu.edu>
Jakob I have not heard about such a problem before. When an HDF5 file is not properly closed, its content may be corrupted. (This will be addressed in the next major release.) There may be two reasons for this: either the file is not closed (which would be an error in the code), or there is a write error (e.g. you run out of disk space). The latter is the major reason for people encountering corrupted HDF5 files. Since you don't see error messages, this is either not the case, or these HDF5 output routines suppress these errors. The thorn SphericalHarmonicDecomp implements its own HDF5 output routines and does not use Cactus. I see that it uses a non-standard way to determine whether the file exists, and that it does not check for errors when writing or closing. I think that HDF5 errors should cause prominent warnings in stdout and stderr (did you check?), and if you don't see these, the writing should have succeeded.The errors I see on stderr are the ones I mentioned in my first mail :
HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5F.c line 1509 in H5Fopen(): unable to open file major: File accessability minor: Unable to open file #001: H5F.c line 1300 in H5F_open(): unable to read superblock major: File accessability minor: Read failed
.... etc.
stdout does not produce any errors or warnings related to this.
You mention checkpointing. Are you experiencing these problems right after recovery, i.e. during the first SphericalHarmonicDecomp HDF5 output afterwards?No, this happened during the first run, not related to recovery.
In this case, did you maybe switch to a new directory where this file doesn't exist? If not, then it may be the non-standard way in which the code determines whether the file already exists, combined with something that may be special about your file system. (The "standard" way operates as follows: open the file as if it existed; if this fails, open it by creating it. The code works differently: it opens the file as binary file. If this fails, the HDF5 file is created; if it succeeds, the file is closed and re-openend as HDF5 file. Maybe the quick closing-then-reopening causes problems?) -erik2012/7/28 Roland Haas <roland.haas@physics.gatech.edu mailto:roland.haas@physics.gatech.edu>
Hello all, >> I have not heard about such a problem before. I believe Nick Taylor at Caltech had similar issues. Bela has since fixed some bugs but had trouble actually committing them (he just saw your emails). I'll grab his changes and commit them.Sounds interesting, I'm looking forward to appying the changes and see if the problem disappears.
Cheers, Jakob
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Dr. Yosef Zlochower Center for Computational Relativity and Gravitation Assistant Professor School of Mathematical Sciences Rochester Institute of Technology 85 Lomb Memorial Drive Rochester, NY 14623
Office:74-2067 Phone: +1 585-475-6103
yosef@astro.rit.edu
CONFIDENTIALITY NOTE: The information transmitted, including attachments, is intended only for the person(s) or entity to which it is addressed and may contain confidential and/or privileged material. Any review, retransmission, dissemination or other use of, or taking of any action in reliance upon this information by persons or entities other than the intended recipient is prohibited. If you received this in error, please contact the sender and destroy any copies of this information.
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 07/28/2012 07:38 AM, Jakob Hansen wrote:
Hi all,
Thanks for fast replys ^^
I just committed a new version of SphericalHarmonicDecomp that checks for IO errors in the hdf5 files. There is a new option SphericalHarmonicDecomp::action_on_hdf5_error which you can set to "abort", which will kill the run on IO errors. Killing the run isn't ideal, but at least you won't have missing timesteps.
2012/7/27 Yosef Zlochower <yosef@astro.rit.edu mailto:yosef@astro.rit.edu>
Hi, SphericalHarmonicDecomp should not be writing output at the same time as a checkpoint. Are you using an NFS mount?No, we're using Lustre filesystem.
I noticed issues with NFS servers becoming unresponsive (due to a large number of blocking io operations) during a checkpoint. Perhaps right after a checkpoint, the server is still too busy.Well, indeed this happens just after checkpointing, however not at every checkpoint and not for all output. I experienced this error twice on two different simulations, once for each simulation. In each case it affected the metric_obs_0_Decomp.h5 right after checkpointing :
Case 1 : Output from ascii_output > gxx.asc : 2.7627600000000001e+02 -7.1882256506489145e-05 1.5746413856874709e-05 2.7640800000000002e+02 -7.1881310717781865e-05 1.5759232641588166e-05 nan 0.0000000000000000e+00 0.0000000000000000e+00 2.7667200000000003e+02 -7.1875929421365471e-05 1.5784711415784759e-05 2.7680400000000003e+02 -7.1871563990008602e-05 1.5797542778149928e-05
In this case there was a checkpoint at time 276.408 : INFO (CarpetIOHDF5): Dumping periodic checkpoint at iteration 268032, simulation time 276.408
Case 2: Output from ascii_output > gxx.asc : 3.2010000000000002e+02 -6.7572912816444132e-05 2.1144268516557760e-05 3.2023200000000003e+02 -6.7570803118733184e-05 2.1156692762978387e-05 nan 0.0000000000000000e+00 0.0000000000000000e+00 3.2049600000000004e+02 -6.7568973330671772e-05 2.1182079214383173e-05 3.2062800000000004e+02 -6.7569118827631724e-05 2.1194791416536239e-05
With a checkpoint at time 320.232 : INFO (CarpetIOHDF5): Dumping periodic checkpoint at iteration 310528, simulation time 320.232
Also, in both cases it only affected the metric_obs_0_Decomp.h5 file, the other detection radius files, _1 and _2, had all data.
<snip>
2012/7/28 Erik Schnetter <schnetter@cct.lsu.edu mailto:schnetter@cct.lsu.edu>
Jakob I have not heard about such a problem before. When an HDF5 file is not properly closed, its content may be corrupted. (This will be addressed in the next major release.) There may be two reasons for this: either the file is not closed (which would be an error in the code), or there is a write error (e.g. you run out of disk space). The latter is the major reason for people encountering corrupted HDF5 files. Since you don't see error messages, this is either not the case, or these HDF5 output routines suppress these errors. The thorn SphericalHarmonicDecomp implements its own HDF5 output routines and does not use Cactus. I see that it uses a non-standard way to determine whether the file exists, and that it does not check for errors when writing or closing. I think that HDF5 errors should cause prominent warnings in stdout and stderr (did you check?), and if you don't see these, the writing should have succeeded.The errors I see on stderr are the ones I mentioned in my first mail :
HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5F.c line 1509 in H5Fopen(): unable to open file major: File accessability minor: Unable to open file #001: H5F.c line 1300 in H5F_open(): unable to read superblock major: File accessability minor: Read failed
.... etc.
stdout does not produce any errors or warnings related to this.
You mention checkpointing. Are you experiencing these problems right after recovery, i.e. during the first SphericalHarmonicDecomp HDF5 output afterwards?No, this happened during the first run, not related to recovery.
In this case, did you maybe switch to a new directory where this file doesn't exist? If not, then it may be the non-standard way in which the code determines whether the file already exists, combined with something that may be special about your file system. (The "standard" way operates as follows: open the file as if it existed; if this fails, open it by creating it. The code works differently: it opens the file as binary file. If this fails, the HDF5 file is created; if it succeeds, the file is closed and re-openend as HDF5 file. Maybe the quick closing-then-reopening causes problems?) -erik2012/7/28 Roland Haas <roland.haas@physics.gatech.edu mailto:roland.haas@physics.gatech.edu>
Hello all, >> I have not heard about such a problem before. I believe Nick Taylor at Caltech had similar issues. Bela has since fixed some bugs but had trouble actually committing them (he just saw your emails). I'll grab his changes and commit them.Sounds interesting, I'm looking forward to appying the changes and see if the problem disappears.
Cheers, Jakob
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 28 Jul 2012, at 13:38, Jakob Hansen wrote:
Hi all,
Thanks for fast replys ^^
2012/7/27 Yosef Zlochower yosef@astro.rit.edu Hi,
SphericalHarmonicDecomp should not be writing output at the same time as a checkpoint. Are you using an NFS mount?
No, we're using Lustre filesystem.
I noticed similar problems on our lustre filesystem several months ago. The symptom was that shortly (a few minutes) after successfully writing a correct HDF5 checkpoint file, some other output would fail. In my case, the output was usually HDF5. The error messages look different though. The only unusual think about my run was that I was doing quite frequent 3D HDF5 output, so I was stressing the system more than usual.
I noticed issues with NFS servers becoming unresponsive (due to a large number of blocking io operations) during a checkpoint. Perhaps right after a checkpoint, the server is still too busy.
Well, indeed this happens just after checkpointing, however not at every checkpoint and not for all output. I experienced this error twice on two different simulations, once for each simulation. In each case it affected the metric_obs_0_Decomp.h5 right after checkpointing :
Case 1 : Output from ascii_output > gxx.asc : 2.7627600000000001e+02 -7.1882256506489145e-05 1.5746413856874709e-05 2.7640800000000002e+02 -7.1881310717781865e-05 1.5759232641588166e-05 nan 0.0000000000000000e+00 0.0000000000000000e+00 2.7667200000000003e+02 -7.1875929421365471e-05 1.5784711415784759e-05 2.7680400000000003e+02 -7.1871563990008602e-05 1.5797542778149928e-05
In this case there was a checkpoint at time 276.408 : INFO (CarpetIOHDF5): Dumping periodic checkpoint at iteration 268032, simulation time 276.408
Case 2: Output from ascii_output > gxx.asc : 3.2010000000000002e+02 -6.7572912816444132e-05 2.1144268516557760e-05 3.2023200000000003e+02 -6.7570803118733184e-05 2.1156692762978387e-05 nan 0.0000000000000000e+00 0.0000000000000000e+00 3.2049600000000004e+02 -6.7568973330671772e-05 2.1182079214383173e-05 3.2062800000000004e+02 -6.7569118827631724e-05 2.1194791416536239e-05
With a checkpoint at time 320.232 : INFO (CarpetIOHDF5): Dumping periodic checkpoint at iteration 310528, simulation time 320.232
Also, in both cases it only affected the metric_obs_0_Decomp.h5 file, the other detection radius files, _1 and _2, had all data.
<snip>
2012/7/28 Erik Schnetter schnetter@cct.lsu.edu Jakob
I have not heard about such a problem before.
When an HDF5 file is not properly closed, its content may be corrupted. (This will be addressed in the next major release.) There may be two reasons for this: either the file is not closed (which would be an error in the code), or there is a write error (e.g. you run out of disk space). The latter is the major reason for people encountering corrupted HDF5 files. Since you don't see error messages, this is either not the case, or these HDF5 output routines suppress these errors.
The thorn SphericalHarmonicDecomp implements its own HDF5 output routines and does not use Cactus. I see that it uses a non-standard way to determine whether the file exists, and that it does not check for errors when writing or closing. I think that HDF5 errors should cause prominent warnings in stdout and stderr (did you check?), and if you don't see these, the writing should have succeeded.
The errors I see on stderr are the ones I mentioned in my first mail :
HDF5-DIAG: Error detected in HDF5 (1.8.5-patch1) thread 0: #000: H5F.c line 1509 in H5Fopen(): unable to open file major: File accessability minor: Unable to open file #001: H5F.c line 1300 in H5F_open(): unable to read superblock major: File accessability minor: Read failed
.... etc.
stdout does not produce any errors or warnings related to this.
You mention checkpointing. Are you experiencing these problems right after recovery, i.e. during the first SphericalHarmonicDecomp HDF5 output afterwards?
No, this happened during the first run, not related to recovery.
In this case, did you maybe switch to a new directory where this file doesn't exist?
If not, then it may be the non-standard way in which the code determines whether the file already exists, combined with something that may be special about your file system.
(The "standard" way operates as follows: open the file as if it existed; if this fails, open it by creating it. The code works differently: it opens the file as binary file. If this fails, the HDF5 file is created; if it succeeds, the file is closed and re-openend as HDF5 file. Maybe the quick closing-then-reopening causes problems?)
-erik
2012/7/28 Roland Haas roland.haas@physics.gatech.edu Hello all,
I have not heard about such a problem before.
I believe Nick Taylor at Caltech had similar issues. Bela has since fixed some bugs but had trouble actually committing them (he just saw your emails). I'll grab his changes and commit them.
Sounds interesting, I'm looking forward to appying the changes and see if the problem disappears.
Cheers, Jakob _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 07/27/2012 06:53 PM, Roland Haas wrote:
Hello all,
I have not heard about such a problem before.
I believe Nick Taylor at Caltech had similar issues. Bela has since fixed some bugs but had trouble actually committing them (he just saw your emails). I'll grab his changes and commit them.
Yours, Roland
Thanks. If you are not ready to commit, can you send a diff between the working version and the current version?
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
Hello Yosef,
Thanks. If you are not ready to commit, can you send a diff between the working version and the current version?
They are up for review on the ET trac:
https://trac.einsteintoolkit.org/ticket/991
although, when looking at them, they should not cause any issues while running the generation code.
As soon as I get a "please apply" I'll commit the changes. The problem of course being that none of the maintainers knows anything about the PITTNullCode so there might be reluctance to review. Yosef, please feel free to review the patches and/or download them.
Yours, Roland
- -- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
On 07/29/2012 04:18 PM, Roland Haas wrote:
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
Hello Yosef,
Thanks. If you are not ready to commit, can you send a diff between the working version and the current version?
They are up for review on the ET trac:
https://trac.einsteintoolkit.org/ticket/991
although, when looking at them, they should not cause any issues while running the generation code.
As soon as I get a "please apply" I'll commit the changes. The problem of course being that none of the maintainers knows anything about the PITTNullCode so there might be reluctance to review. Yosef, please feel free to review the patches and/or download them.
Yours, Roland
Are all of them there? I didn't see a patch for the SphericalHarmonicDecomp, Recomp codes?
My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net. -----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.12 (GNU/Linux) Comment: Using GnuPG with Mozilla - http://enigmail.mozdev.org/
iEYEARECAAYFAlAVmqQACgkQTiFSTN7SboXpCgCgyv7gA95OqiMKm0/5mRTymzN8 utkAn3kMCNX9V8x2YSN1oS6rpnXvE0if =FAs/ -----END PGP SIGNATURE-----
Hello Yosef,
Are all of them there? I didn't see a patch for the SphericalHarmonicDecomp, Recomp codes?
They are all I have. There were patches to Christian Reisswig's SphericalHarmonicReconASCII in incoming. Since then the thorn has changed names to SphericalHarmonicReconGen. Unfortunately I don't think any of these will help with a run that produces faulty output from Cactus (since the input to CCE for Caltech came from SpEC).
Yours, Roland
Hello Yosef,
Are all of them there? I didn't see a patch for the SphericalHarmonicDecomp, Recomp codes?
They are all I have. There were patches to Christian Reisswig's SphericalHarmonicReconASCII in incoming. Since then the thorn has changed names to SphericalHarmonicReconGen. Unfortunately I don't think any of these will help with a run that produces faulty output from Cactus (since the input to CCE for Caltech came from SpEC).
I believe the most important patch is the one that fixes a problem with the start-up algorithm at the worldtube. None of the patches affects the SphericalHarmonicRecon/Decomp thorns.
Indeed, at Caltech, we are using SphericalHarmonicReconGen (formerly SphericalHarmonicReconASCII) which can read the data from the SpEC code. This thorn should be located in the incoming directory and I hope it can make it to the next release.
So I believe your problems will not be fixed by the changes listed in ticket 991.
cheers, Christian
On 07/30/2012 02:22 AM, Christian Reisswig wrote:
Hello Yosef,
Are all of them there? I didn't see a patch for the SphericalHarmonicDecomp, Recomp codes?
They are all I have. There were patches to Christian Reisswig's SphericalHarmonicReconASCII in incoming. Since then the thorn has changed names to SphericalHarmonicReconGen. Unfortunately I don't think any of these will help with a run that produces faulty output from Cactus (since the input to CCE for Caltech came from SpEC).
I believe the most important patch is the one that fixes a problem with the start-up algorithm at the worldtube.
Does this patch fix the O(h) error in the waveform in the CCE paper?
None of the patches affects the SphericalHarmonicRecon/Decomp thorns.
Indeed, at Caltech, we are using SphericalHarmonicReconGen (formerly SphericalHarmonicReconASCII) which can read the data from the SpEC code. This thorn should be located in the incoming directory and I hope it can make it to the next release.
So I believe your problems will not be fixed by the changes listed in ticket 991.
cheers, Christian _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 07/30/2012 02:22 AM, Christian Reisswig wrote:
Hello Yosef,
Are all of them there? I didn't see a patch for the SphericalHarmonicDecomp, Recomp codes?
They are all I have. There were patches to Christian Reisswig's SphericalHarmonicReconASCII in incoming. Since then the thorn has changed names to SphericalHarmonicReconGen. Unfortunately I don't think any of these will help with a run that produces faulty output from Cactus (since the input to CCE for Caltech came from SpEC).
I believe the most important patch is the one that fixes a problem with the start-up algorithm at the worldtube. None of the patches affects the SphericalHarmonicRecon/Decomp thorns.
I looked through the patches and I don't see where the start-up algorithm is modified. Also there are these two changes that I am not sure about. The first changes the algorithm in the middle of a run. The second changes the meaning of the "time" in the output file. My guess is that a typical user wouldn't want to do either of these. Someone doing a test, could modify the param.ccl themselves.
-BOOLEAN first_order_scheme "should angular derviatives be reduced to first order?" +BOOLEAN first_order_scheme "should angular derviatives be reduced to first order?" STEERABLE=ALWAYS
-BOOLEAN interp_to_constant_uBondi "Interpolate quantities at Scri to constant Bondi time" +BOOLEAN interp_to_constant_uBondi "Interpolate quantities at Scri to constant Bondi time" STEERABLE=ALWAYS
Indeed, at Caltech, we are using SphericalHarmonicReconGen (formerly SphericalHarmonicReconASCII) which can read the data from the SpEC code. This thorn should be located in the incoming directory and I hope it can make it to the next release.
So I believe your problems will not be fixed by the changes listed in ticket 991.
cheers, Christian _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Yosef, and all
'start-up' here refers to the way the null parallelogram algorithm (marching out along a characteristic slice and starting from the inner boundary), gets its integration constants set by the boundary data. This is not a t=0 issue, but rather a \lambda=0 issue in CCE language. I believe it is absolutely essential for the correctness of the code.
As far as changing parameters from non-steerable to steerable goes, I also believe they are useful as one can legitimately want to start a new diagnostic for a run when restarted from a particular checkpoint. Or even from the middle of a run, checkpointed or not. Those are, indeed, independent from the 'start-up' patch.
I do understand that the Einstein Toolkit maintenance team may not be familiar with the Pitt code. Please regard the 'start-up' patch as an amendment to the 1st release of the code, by those same people who worked on it before releasing it. If you need any details of why have I changed things in the particular way I have, I will be glad to explain.
Bela.
On Mon, Jul 30, 2012 at 6:12 AM, Yosef Zlochower yosef@astro.rit.edu wrote:
On 07/30/2012 02:22 AM, Christian Reisswig wrote:
Hello Yosef,
Are all of them there? I didn't see a patch for the SphericalHarmonicDecomp, Recomp codes?
They are all I have. There were patches to Christian Reisswig's SphericalHarmonicReconASCII in incoming. Since then the thorn has changed names to SphericalHarmonicReconGen. Unfortunately I don't think any of these will help with a run that produces faulty output from Cactus (since the input to CCE for Caltech came from SpEC).
I believe the most important patch is the one that fixes a problem with the start-up algorithm at the worldtube. None of the patches affects the SphericalHarmonicRecon/Decomp thorns.
I looked through the patches and I don't see where the start-up algorithm is modified. Also there are these two changes that I am not sure about. The first changes the algorithm in the middle of a run. The second changes the meaning of the "time" in the output file. My guess is that a typical user wouldn't want to do either of these. Someone doing a test, could modify the param.ccl themselves.
-BOOLEAN first_order_scheme "should angular derviatives be reduced to first order?" +BOOLEAN first_order_scheme "should angular derviatives be reduced to first order?" STEERABLE=ALWAYS
-BOOLEAN interp_to_constant_uBondi "Interpolate quantities at Scri to constant Bondi time" +BOOLEAN interp_to_constant_uBondi "Interpolate quantities at Scri to constant Bondi time" STEERABLE=ALWAYS
Indeed, at Caltech, we are using SphericalHarmonicReconGen (formerly SphericalHarmonicReconASCII) which can read the data from the SpEC code. This thorn should be located in the incoming directory and I hope it can make it to the next release.
So I believe your problems will not be fixed by the changes listed in ticket 991.
cheers, Christian _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Dr. Yosef Zlochower Center for Computational Relativity and Gravitation Assistant Professor School of Mathematical Sciences Rochester Institute of Technology 85 Lomb Memorial Drive Rochester, NY 14623
Office:74-2067 Phone: +1 585-475-6103
yosef@astro.rit.edu
CONFIDENTIALITY NOTE: The information transmitted, including attachments, is intended only for the person(s) or entity to which it is addressed and may contain confidential and/or privileged material. Any review, retransmission, dissemination or other use of, or taking of any action in reliance upon this information by persons or entities other than the intended recipient is prohibited. If you received this in error, please contact the sender and destroy any copies of this information. _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Bela,
Does this patch fix the O(h) error we saw in the CCE paper? I think this patch was missing. Can you submit it?
On 07/30/2012 11:47 AM, Bela Szilagyi wrote:
Yosef, and all
'start-up' here refers to the way the null parallelogram algorithm (marching out along a characteristic slice and starting from the inner boundary), gets its integration constants set by the boundary data. This is not a t=0 issue, but rather a \lambda=0 issue in CCE language. I believe it is absolutely essential for the correctness of the code.
As far as changing parameters from non-steerable to steerable goes, I also believe they are useful as one can legitimately want to start a new diagnostic for a run when restarted from a particular checkpoint. Or even from the middle of a run, checkpointed or not. Those are, indeed, independent from the 'start-up' patch.
I do understand that the Einstein Toolkit maintenance team may not be familiar with the Pitt code. Please regard the 'start-up' patch as an amendment to the 1st release of the code, by those same people who worked on it before releasing it. If you need any details of why have I changed things in the particular way I have, I will be glad to explain.
Bela.
On Mon, Jul 30, 2012 at 6:12 AM, Yosef Zlochoweryosef@astro.rit.edu wrote:
On 07/30/2012 02:22 AM, Christian Reisswig wrote:
Hello Yosef,
Are all of them there? I didn't see a patch for the SphericalHarmonicDecomp, Recomp codes?
They are all I have. There were patches to Christian Reisswig's SphericalHarmonicReconASCII in incoming. Since then the thorn has changed names to SphericalHarmonicReconGen. Unfortunately I don't think any of these will help with a run that produces faulty output from Cactus (since the input to CCE for Caltech came from SpEC).
I believe the most important patch is the one that fixes a problem with the start-up algorithm at the worldtube. None of the patches affects the SphericalHarmonicRecon/Decomp thorns.
I looked through the patches and I don't see where the start-up algorithm is modified. Also there are these two changes that I am not sure about. The first changes the algorithm in the middle of a run. The second changes the meaning of the "time" in the output file. My guess is that a typical user wouldn't want to do either of these. Someone doing a test, could modify the param.ccl themselves.
-BOOLEAN first_order_scheme "should angular derviatives be reduced to first order?" +BOOLEAN first_order_scheme "should angular derviatives be reduced to first order?" STEERABLE=ALWAYS
-BOOLEAN interp_to_constant_uBondi "Interpolate quantities at Scri to constant Bondi time" +BOOLEAN interp_to_constant_uBondi "Interpolate quantities at Scri to constant Bondi time" STEERABLE=ALWAYS
Indeed, at Caltech, we are using SphericalHarmonicReconGen (formerly SphericalHarmonicReconASCII) which can read the data from the SpEC code. This thorn should be located in the incoming directory and I hope it can make it to the next release.
So I believe your problems will not be fixed by the changes listed in ticket 991.
cheers, Christian _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
The 1st order accuracy is yet another story. Part of it may be induced by the news code itself but understanding and fixing this will need more effort.
Sent from my iPhone
On Jul 30, 2012, at 8:51 AM, Yosef Zlochower yosef@astro.rit.edu wrote:
Bela,
Does this patch fix the O(h) error we saw in the CCE paper? I think this patch was missing. Can you submit it?
On 07/30/2012 11:47 AM, Bela Szilagyi wrote:
Yosef, and all
'start-up' here refers to the way the null parallelogram algorithm (marching out along a characteristic slice and starting from the inner boundary), gets its integration constants set by the boundary data. This is not a t=0 issue, but rather a \lambda=0 issue in CCE language. I believe it is absolutely essential for the correctness of the code.
As far as changing parameters from non-steerable to steerable goes, I also believe they are useful as one can legitimately want to start a new diagnostic for a run when restarted from a particular checkpoint. Or even from the middle of a run, checkpointed or not. Those are, indeed, independent from the 'start-up' patch.
I do understand that the Einstein Toolkit maintenance team may not be familiar with the Pitt code. Please regard the 'start-up' patch as an amendment to the 1st release of the code, by those same people who worked on it before releasing it. If you need any details of why have I changed things in the particular way I have, I will be glad to explain.
Bela.
On Mon, Jul 30, 2012 at 6:12 AM, Yosef Zlochoweryosef@astro.rit.edu wrote:
On 07/30/2012 02:22 AM, Christian Reisswig wrote:
Hello Yosef,
> Are all of them there? I didn't see a patch for the > SphericalHarmonicDecomp, Recomp > codes?
They are all I have. There were patches to Christian Reisswig's SphericalHarmonicReconASCII in incoming. Since then the thorn has changed names to SphericalHarmonicReconGen. Unfortunately I don't think any of these will help with a run that produces faulty output from Cactus (since the input to CCE for Caltech came from SpEC).
I believe the most important patch is the one that fixes a problem with the start-up algorithm at the worldtube. None of the patches affects the SphericalHarmonicRecon/Decomp thorns.
I looked through the patches and I don't see where the start-up algorithm is modified. Also there are these two changes that I am not sure about. The first changes the algorithm in the middle of a run. The second changes the meaning of the "time" in the output file. My guess is that a typical user wouldn't want to do either of these. Someone doing a test, could modify the param.ccl themselves.
-BOOLEAN first_order_scheme "should angular derviatives be reduced to first order?" +BOOLEAN first_order_scheme "should angular derviatives be reduced to first order?" STEERABLE=ALWAYS
-BOOLEAN interp_to_constant_uBondi "Interpolate quantities at Scri to constant Bondi time" +BOOLEAN interp_to_constant_uBondi "Interpolate quantities at Scri to constant Bondi time" STEERABLE=ALWAYS
Indeed, at Caltech, we are using SphericalHarmonicReconGen (formerly SphericalHarmonicReconASCII) which can read the data from the SpEC code. This thorn should be located in the incoming directory and I hope it can make it to the next release.
So I believe your problems will not be fixed by the changes listed in ticket 991.
cheers, Christian _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 07/30/2012 11:47 AM, Bela Szilagyi wrote:
Yosef, and all
'start-up' here refers to the way the null parallelogram algorithm (marching out along a characteristic slice and starting from the inner boundary), gets its integration constants set by the boundary data. This is not a t=0 issue, but rather a \lambda=0 issue in CCE language. I believe it is absolutely essential for the correctness of the code.
As far as changing parameters from non-steerable to steerable goes, I also believe they are useful as one can legitimately want to start a new diagnostic for a run when restarted from a particular checkpoint. Or even from the middle of a run, checkpointed or not. Those are, indeed, independent from the 'start-up' patch.
I do understand that the Einstein Toolkit maintenance team may not be familiar with the Pitt code. Please regard the 'start-up' patch as an amendment to the 1st release of the code, by those same people who worked on it before releasing it. If you need any details of why have I changed things in the particular way I have, I will be glad to explain.
Bela et al, It would probably be a good idea to have some kind of informative commit message. Since you checked that these changes are correct and necessary, can you commit them? I'll then update the testsuite, if needed.
Bela.
Yosef,
I have made several attempts to commit the changes, over the course of the past couple of months. Somehow every time I figure out how to re-download the code from the latest incarnation of the repository, I figure I no longer have write access or no longer know my password or some combination of the above. This simply because I am not an active Cactus user.
I had asked Roland to just push these changes as they are essential for the correctness of the code. If you rely on me finding the time to figure out how to (re-)gain checking access to these thorns, this will delay the application of the patches by further months.
The start-up related change could be logged as
Improve the start-up algorithm of the characteristic marching scheme from the inner boundary towards scri+.
All other changes are related to either improving IO (will correctly truncate/create diagnostic files) or allowing for flexible tuning of IO between checkpoints/restarts (by changing the parameters into steerable ones).
I'd like to ask that someone with working write access to the arrangement apply the patches.
Bela.
On Mon, Jul 30, 2012 at 12:05 PM, Yosef Zlochower yosef@astro.rit.edu wrote:
On 07/30/2012 11:47 AM, Bela Szilagyi wrote:
Yosef, and all
'start-up' here refers to the way the null parallelogram algorithm (marching out along a characteristic slice and starting from the inner boundary), gets its integration constants set by the boundary data. This is not a t=0 issue, but rather a \lambda=0 issue in CCE language. I believe it is absolutely essential for the correctness of the code.
As far as changing parameters from non-steerable to steerable goes, I also believe they are useful as one can legitimately want to start a new diagnostic for a run when restarted from a particular checkpoint. Or even from the middle of a run, checkpointed or not. Those are, indeed, independent from the 'start-up' patch.
I do understand that the Einstein Toolkit maintenance team may not be familiar with the Pitt code. Please regard the 'start-up' patch as an amendment to the 1st release of the code, by those same people who worked on it before releasing it. If you need any details of why have I changed things in the particular way I have, I will be glad to explain.
Bela et al, It would probably be a good idea to have some kind of informative commit message. Since you checked that these changes are correct and necessary, can you commit them? I'll then update the testsuite, if needed.
Bela.
I updated the trac ticket. btw, these changes do not affect the testsuite. I'm actually surprised by this. Does the bugfix only affect the code under certain circumstances?
On 07/31/2012 02:49 PM, Bela Szilagyi wrote:
Yosef,
I have made several attempts to commit the changes, over the course of the past couple of months. Somehow every time I figure out how to re-download the code from the latest incarnation of the repository, I figure I no longer have write access or no longer know my password or some combination of the above. This simply because I am not an active Cactus user.
I had asked Roland to just push these changes as they are essential for the correctness of the code. If you rely on me finding the time to figure out how to (re-)gain checking access to these thorns, this will delay the application of the patches by further months.
The start-up related change could be logged as
Improve the start-up algorithm of the characteristic marching scheme from the inner boundary towards scri+.
All other changes are related to either improving IO (will correctly truncate/create diagnostic files) or allowing for flexible tuning of IO between checkpoints/restarts (by changing the parameters into steerable ones).
I'd like to ask that someone with working write access to the arrangement apply the patches.
Bela.
On Mon, Jul 30, 2012 at 12:05 PM, Yosef Zlochoweryosef@astro.rit.edu wrote:
On 07/30/2012 11:47 AM, Bela Szilagyi wrote:
Yosef, and all
'start-up' here refers to the way the null parallelogram algorithm (marching out along a characteristic slice and starting from the inner boundary), gets its integration constants set by the boundary data. This is not a t=0 issue, but rather a \lambda=0 issue in CCE language. I believe it is absolutely essential for the correctness of the code.
As far as changing parameters from non-steerable to steerable goes, I also believe they are useful as one can legitimately want to start a new diagnostic for a run when restarted from a particular checkpoint. Or even from the middle of a run, checkpointed or not. Those are, indeed, independent from the 'start-up' patch.
I do understand that the Einstein Toolkit maintenance team may not be familiar with the Pitt code. Please regard the 'start-up' patch as an amendment to the 1st release of the code, by those same people who worked on it before releasing it. If you need any details of why have I changed things in the particular way I have, I will be glad to explain.
Bela et al, It would probably be a good idea to have some kind of informative commit message. Since you checked that these changes are correct and necessary, can you commit them? I'll then update the testsuite, if needed.
Bela.
Yosef,
in order for the bug to show, the world tube had to be at a distance A * dt *1/4 from the nearest Bondi-frame evolution point,where 'A' is some geometric factor, dependent on the metric. What we had not realized the last time this code was looked at, is that we were computing one of the corner points of the null parallelogram by an interpolation stencil that divided by a quantity that can vanish when the world tube was at this 'wrong' distance.
I rewrote the interpolation (and redefined the 'inner' end of the null parallelogram) such that this type of condition cannot occur anymore.
What we noticed was that in certain runs of ours there would be spikes in the waveform at scri+. We managed to correlate this problem to the world-tube having moved across this point where the interpolation became singular. After the fix the spikes went away.
Bela.
On Tue, Jul 31, 2012 at 12:06 PM, Yosef Zlochower yosef@astro.rit.edu wrote:
I updated the trac ticket. btw, these changes do not affect the testsuite. I'm actually surprised by this. Does the bugfix only affect the code under certain circumstances?
On 07/31/2012 02:49 PM, Bela Szilagyi wrote:
Yosef,
I have made several attempts to commit the changes, over the course of the past couple of months. Somehow every time I figure out how to re-download the code from the latest incarnation of the repository, I figure I no longer have write access or no longer know my password or some combination of the above. This simply because I am not an active Cactus user.
I had asked Roland to just push these changes as they are essential for the correctness of the code. If you rely on me finding the time to figure out how to (re-)gain checking access to these thorns, this will delay the application of the patches by further months.
The start-up related change could be logged as
Improve the start-up algorithm of the characteristic marching scheme from the inner boundary towards scri+.
All other changes are related to either improving IO (will correctly truncate/create diagnostic files) or allowing for flexible tuning of IO between checkpoints/restarts (by changing the parameters into steerable ones).
I'd like to ask that someone with working write access to the arrangement apply the patches.
Bela.
On Mon, Jul 30, 2012 at 12:05 PM, Yosef Zlochoweryosef@astro.rit.edu wrote:
On 07/30/2012 11:47 AM, Bela Szilagyi wrote:
Yosef, and all
'start-up' here refers to the way the null parallelogram algorithm (marching out along a characteristic slice and starting from the inner boundary), gets its integration constants set by the boundary data. This is not a t=0 issue, but rather a \lambda=0 issue in CCE language. I believe it is absolutely essential for the correctness of the code.
As far as changing parameters from non-steerable to steerable goes, I also believe they are useful as one can legitimately want to start a new diagnostic for a run when restarted from a particular checkpoint. Or even from the middle of a run, checkpointed or not. Those are, indeed, independent from the 'start-up' patch.
I do understand that the Einstein Toolkit maintenance team may not be familiar with the Pitt code. Please regard the 'start-up' patch as an amendment to the 1st release of the code, by those same people who worked on it before releasing it. If you need any details of why have I changed things in the particular way I have, I will be glad to explain.
Bela et al, It would probably be a good idea to have some kind of informative commit message. Since you checked that these changes are correct and necessary, can you commit them? I'll then update the testsuite, if needed.
Bela.
-- Dr. Yosef Zlochower Center for Computational Relativity and Gravitation Assistant Professor School of Mathematical Sciences Rochester Institute of Technology 85 Lomb Memorial Drive Rochester, NY 14623
Office:74-2067 Phone: +1 585-475-6103
yosef@astro.rit.edu
CONFIDENTIALITY NOTE: The information transmitted, including attachments, is intended only for the person(s) or entity to which it is addressed and may contain confidential and/or privileged material. Any review, retransmission, dissemination or other use of, or taking of any action in reliance upon this information by persons or entities other than the intended recipient is prohibited. If you received this in error, please contact the sender and destroy any copies of this information.
Hello Yosef,
I looked through the patches and I don't see where the start-up algorithm is modified. Also there are these two changes that I am not sure about. The first changes the algorithm in the middle of a run.
It seems as if the patches were mixed up somehow (though just how that happened is a mystery to me, given that it is just a svn diff inside of a bash for loop). Anyhow. I have updated the patches on the trac ticket and now one (NullEvolve) actually has non-trivial code changes. There are now to 1 and 2 byte patch files since I cannot seem to remove attachments, only replace them.
The second changes the meaning of the "time" in the output file. My guess is that a typical user wouldn't want to do either of these.
Can you point out which change this is, please?
-BOOLEAN first_order_scheme "should angular derviatives be reduced to first order?" +BOOLEAN first_order_scheme "should angular derviatives be reduced to first order?" STEERABLE=ALWAYS
-BOOLEAN interp_to_constant_uBondi "Interpolate quantities at Scri to constant Bondi time" +BOOLEAN interp_to_constant_uBondi "Interpolate quantities at Scri to constant Bondi time" STEERABLE=ALWAYS
These are the ones that allow changing the algorithm in the middle of arun? Or is the first_order_scheme the one that changes algorithm and the Bondi time one the one that changes output time?
I usually promote the viewpoint that parameters should be as steerable as possible without breaking the code, even if a means that a user can ruin their data by unwise parameter changes (rather than just unwise parameter settings). maybe being able to steer this parameter at least during recovery is useful for some debugging runs?
Yours, Roland
On 07/30/2012 11:54 AM, Roland Haas wrote:
Hello Yosef,
I looked through the patches and I don't see where the start-up algorithm is modified. Also there are these two changes that I am not sure about. The first changes the algorithm in the middle of a run.
It seems as if the patches were mixed up somehow (though just how that happened is a mystery to me, given that it is just a svn diff inside of a bash for loop). Anyhow. I have updated the patches on the trac ticket and now one (NullEvolve) actually has non-trivial code changes. There are now to 1 and 2 byte patch files since I cannot seem to remove attachments, only replace them.
The second changes the meaning of the "time" in the output file. My guess is that a typical user wouldn't want to do either of these.
Can you point out which change this is, please?
This would happen if you steer the parameter "interp_to_constant_uBond" during a run. If this is set to no, the output files contain spherical harmonics of the News (etc) at fixed code time (and the label may be off by 1/2 dt). If it's switched to yes, then the same files would contain the harmonics calculated on a fixed bondi time.
-BOOLEAN first_order_scheme "should angular derviatives be reduced to first order?" +BOOLEAN first_order_scheme "should angular derviatives be reduced to first order?" STEERABLE=ALWAYS
-BOOLEAN interp_to_constant_uBondi "Interpolate quantities at Scri to constant Bondi time" +BOOLEAN interp_to_constant_uBondi "Interpolate quantities at Scri to constant Bondi time" STEERABLE=ALWAYS
These are the ones that allow changing the algorithm in the middle of arun? Or is the first_order_scheme the one that changes algorithm and the Bondi time one the one that changes output time?
I usually promote the viewpoint that parameters should be as steerable as possible without breaking the code, even if a means that a user can ruin their data by unwise parameter changes (rather than just unwise parameter settings). maybe being able to steer this parameter at least during recovery is useful for some debugging runs?
Yours, Roland
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
users@lists.einsteintoolkit.org