Hi,
I have three patches to Multipole:
===
1. Convert to C++ strings
2. Use IO::out_dir parameter by sharing from implementation rather than via CCTK_ParameterGet
3. Add HDF5 output support
One HDF5 file per variable, and one extensible dataset per radius per mode. HDF5 is required only optionally, so this thorn can be compiled without it
===
The advantage of the HDF5 output is that you get only one file per variable, instead of the hundreds that you get with the previous one file per variable per mode per radius format. Instead of using ASCII output, which is unindexed and very inefficient to parse, HDF5 output is structured and a single dataset can be easily read. Analysis tools such as Python and Mathematica support reading HDF5, and the h5dump or h5totxt utility (in the h5utils package) can be used to convert the file to text for any tool which can't read HDF5.
Two new parameters are introduced: output_ascii (defaults to yes), and output_hdf5 (defaults to no).
Example output (modes are the same for all t in this test case; columns are t, re(var), im(var)):
$ h5ls mp_harmonic.h5 l0_m0_r8.00 Dataset {11/Inf, 3} l1_m-1_r8.00 Dataset {11/Inf, 3} l1_m0_r8.00 Dataset {11/Inf, 3} l1_m1_r8.00 Dataset {11/Inf, 3} l2_m-1_r8.00 Dataset {11/Inf, 3} l2_m-2_r8.00 Dataset {11/Inf, 3} l2_m0_r8.00 Dataset {11/Inf, 3} l2_m1_r8.00 Dataset {11/Inf, 3} l2_m2_r8.00 Dataset {11/Inf, 3}
$ h5totxt -d l2_m2_r8.00 mp_harmonic.h5 0,1.009144071427338,1.00100865293711e-06 1,1.009144071427338,1.00100865293711e-06 2,1.009144071427338,1.00100865293711e-06 3,1.009144071427338,1.00100865293711e-06 4,1.009144071427338,1.00100865293711e-06 5,1.009144071427338,1.00100865293711e-06 6,1.009144071427338,1.00100865293711e-06 7,1.009144071427338,1.00100865293711e-06 8,1.009144071427338,1.00100865293711e-06 9,1.009144071427338,1.00100865293711e-06 10,1.009144071427338,1.00100865293711e-06
$ h5dump -d l2_m2_r8.00 mp_harmonic.h5 HDF5 "mp_harmonic.h5" { DATASET "l2_m2_r8.00" { DATATYPE H5T_IEEE_F64LE DATASPACE SIMPLE { ( 11, 3 ) / ( H5S_UNLIMITED, 3 ) } DATA { (0,0): 0, 1.00914, 1.00101e-06, (1,0): 1, 1.00914, 1.00101e-06, (2,0): 2, 1.00914, 1.00101e-06, (3,0): 3, 1.00914, 1.00101e-06, (4,0): 4, 1.00914, 1.00101e-06, (5,0): 5, 1.00914, 1.00101e-06, (6,0): 6, 1.00914, 1.00101e-06, (7,0): 7, 1.00914, 1.00101e-06, (8,0): 8, 1.00914, 1.00101e-06, (9,0): 9, 1.00914, 1.00101e-06, (10,0): 10, 1.00914, 1.00101e-06 } } }
Testsuites pass. OK to commit?
PS: The optional capability support depends on a recent patch to the flesh which has not been applied yet. I won't commit these three patches unless that patch is accepted and committed.
On 9 Dec 2010, at 11:07, Eloisa Bentivegna wrote:
On Dec 9, 2010, at 9:43 AM, Ian Hinder wrote:
Hi,
I have three patches to Multipole: ... Testsuites pass. OK to commit?
One question: does the variable name appear anywhere in the filename or metadata?
Yes, the filename is mp_<varname>.h5. This variable name is the same name that is used for the filename in the ascii output, which defaults to the real part that you specify in the "variables" parameter, but can be overridden by specifying a "name" option for the variable (since it doesn't make sense for the output file to be called Psi4r when that is only the real part):
Multipole::variables = "Multipole::harmonic_re{sw=-2 cmplx='Multipole::harmonic_im' name='harmonic'}"
Hi,
I support all three ideas (but didn't look at the patched yet).
On Thu, Dec 09, 2010 at 09:43:04AM +0100, Ian Hinder wrote:
Two new parameters are introduced: output_ascii (defaults to yes), and output_hdf5 (defaults to no).
Do you think that most people would actually use the hdf5 version? Could it be made the default output if HDF5 is found, disabling ascii in that case (unless requested by setting this explicitly by the parameter)?
That would probably mean that both parameters also accept something like "auto" which has exactly that meaning: hdf5 if found, otherwise ascii.
Frank
On 9 Dec 2010, at 15:54, Frank Loeffler wrote:
Hi,
I support all three ideas (but didn't look at the patched yet).
On Thu, Dec 09, 2010 at 09:43:04AM +0100, Ian Hinder wrote:
Two new parameters are introduced: output_ascii (defaults to yes), and output_hdf5 (defaults to no).
Do you think that most people would actually use the hdf5 version?
I expect that people have very high inertia. However, on production filesystems, having a large number of output files is frowned upon, and makes things quite slow, so I would encourage people to use HDF5 output whenever possible.
Could it be made the default output if HDF5 is found, disabling ascii in that case (unless requested by setting this explicitly by the parameter)?
I think this would be surprising to some people. Imagine that you were happily using ASCII output in Multipole, and one day decided to include the HDF5 thorn in your thornlist because of something completely unrelated. The next time you ran, you would find that there was no ASCII output for Multipole, and all your analysis scripts would not be able to deal with the HDF5 data. I would adhere to the principle of least surprise.
If we want to encourage people to use HDF5 output, we could enable HDF5 by default, but I would not like to disable ASCII by default just yet. I don't want to enable HDF5 by default just yet, as I don't know what the performance impact of it will be. Especially on clusters with slow filesystems (I'm looking at you, Kraken). We could enable it by default after it has seen some more testing.
That would probably mean that both parameters also accept something like "auto" which has exactly that meaning: hdf5 if found, otherwise ascii.
On Thu, Dec 09, 2010 at 04:43:53PM +0100, Ian Hinder wrote:
Do you think that most people would actually use the hdf5 version?
I expect that people have very high inertia. However, on production filesystems, having a large number of output files is frowned upon, and makes things quite slow, so I would encourage people to use HDF5 output whenever possible.
That is why I asked. I would expect most people to use the hdf5 version now.
I think this would be surprising to some people. Imagine that you were happily using ASCII output in Multipole, and one day decided to include the HDF5 thorn in your thornlist because of something completely unrelated. The next time you ran, you would find that there was no ASCII output for Multipole, and all your analysis scripts would not be able to deal with the HDF5 data. I would adhere to the principle of least surprise.
I agree that this could happen. What about if the thorn would provide a utility (script) which uses h5dump to create the ascii-output from the hdf5 one?
If we want to encourage people to use HDF5 output, we could enable HDF5 by default, but I would not like to disable ASCII by default just yet. I don't want to enable HDF5 by default just yet, as I don't know what the performance impact of it will be. Especially on clusters with slow filesystems (I'm looking at you, Kraken). We could enable it by default after it has seen some more testing.
Yes, we could do that. However, production runs wouldn't all see the new version right now anyway (most of them will probably use the released version, right).
Frank
On 9 Dec 2010, at 17:46, Frank Loeffler wrote:
On Thu, Dec 09, 2010 at 04:43:53PM +0100, Ian Hinder wrote:
Do you think that most people would actually use the hdf5 version?
I expect that people have very high inertia. However, on production filesystems, having a large number of output files is frowned upon, and makes things quite slow, so I would encourage people to use HDF5 output whenever possible.
That is why I asked. I would expect most people to use the hdf5 version now.
I think this would be surprising to some people. Imagine that you were happily using ASCII output in Multipole, and one day decided to include the HDF5 thorn in your thornlist because of something completely unrelated. The next time you ran, you would find that there was no ASCII output for Multipole, and all your analysis scripts would not be able to deal with the HDF5 data. I would adhere to the principle of least surprise.
I agree that this could happen. What about if the thorn would provide a utility (script) which uses h5dump to create the ascii-output from the hdf5 one?
This is a good idea in any case. But I don't think that activating or deactivating a thorn should change the output behaviour of another thorn - it is too unexpected.
If we want to encourage people to use HDF5 output, we could enable HDF5 by default, but I would not like to disable ASCII by default just yet. I don't want to enable HDF5 by default just yet, as I don't know what the performance impact of it will be. Especially on clusters with slow filesystems (I'm looking at you, Kraken). We could enable it by default after it has seen some more testing.
Yes, we could do that. However, production runs wouldn't all see the new version right now anyway (most of them will probably use the released version, right).
Do people use the released version for production runs? I don't. But then, I also don't routinely update my main Cactus tree. There is little enough development going on in critical areas that this doesn't seem important.
Hi,
On Fri, Dec 10, 2010 at 11:42:42AM +0100, Ian Hinder wrote:
This is a good idea in any case. But I don't think that activating or deactivating a thorn should change the output behaviour of another thorn - it is too unexpected.
Isn't that what would be introduced with the optional dependency on HDF5? One could see at this way: by default hdf5 is used, but there is also a fallback to ascii in case someone doesn't have hdf5 for whatever reason - so that the thorn is still usable. No parameter file would have to be changed for example. The simulation would still run.
Do people use the released version for production runs? I don't. But then, I also don't routinely update my main Cactus tree. There is little enough development going on in critical areas that this doesn't seem important.
Right. Therefore I think it would be ok to change the default behavior in the development version, as long as we don't forget to mention this in the release notes.
Frank
Let's add the new feature first, and not try to do too many things at once. For example, we haven't even tried the HDF5 output yet. There is no urgency to switch defaults now.
Let's apply the patch now (or once it follows the Cactus conventions for truncating output files), and then see what happens. Let's not hold up the patch with this discussion.
-erik
On Fri, Dec 10, 2010 at 11:09 AM, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi,
On Fri, Dec 10, 2010 at 11:42:42AM +0100, Ian Hinder wrote:
This is a good idea in any case. But I don't think that activating or deactivating a thorn should change the output behaviour of another thorn - it is too unexpected.
Isn't that what would be introduced with the optional dependency on HDF5? One could see at this way: by default hdf5 is used, but there is also a fallback to ascii in case someone doesn't have hdf5 for whatever reason - so that the thorn is still usable. No parameter file would have to be changed for example. The simulation would still run.
Do people use the released version for production runs? I don't. But then, I also don't routinely update my main Cactus tree. There is little enough development going on in critical areas that this doesn't seem important.
Right. Therefore I think it would be ok to change the default behavior in the development version, as long as we don't forget to mention this in the release notes.
Frank
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 17 Dec 2010, at 22:09, Erik Schnetter wrote:
Let's add the new feature first, and not try to do too many things at once. For example, we haven't even tried the HDF5 output yet. There is no urgency to switch defaults now.
Let's apply the patch now (or once it follows the Cactus conventions for truncating output files), and then see what happens. Let's not hold up the patch with this discussion.
I agree. I have made the change requested and the modified patch series is attached (only the last patch has changed). OK to commit once the flesh patch has been committed?
-erik
On Fri, Dec 10, 2010 at 11:09 AM, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi,
On Fri, Dec 10, 2010 at 11:42:42AM +0100, Ian Hinder wrote:
This is a good idea in any case. But I don't think that activating or deactivating a thorn should change the output behaviour of another thorn - it is too unexpected.
Isn't that what would be introduced with the optional dependency on HDF5? One could see at this way: by default hdf5 is used, but there is also a fallback to ascii in case someone doesn't have hdf5 for whatever reason - so that the thorn is still usable. No parameter file would have to be changed for example. The simulation would still run.
Do people use the released version for production runs? I don't. But then, I also don't routinely update my main Cactus tree. There is little enough development going on in critical areas that this doesn't seem important.
Right. Therefore I think it would be ok to change the default behavior in the development version, as long as we don't forget to mention this in the release notes.
On Thu, Dec 9, 2010 at 8:43 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
Hi,
I have three patches to Multipole:
===
Convert to C++ strings
Use IO::out_dir parameter by sharing from implementation rather than via CCTK_ParameterGet
Add HDF5 output support
The standard behaviour for Cactus output files is that files are deleted/truncated unless this is a recovery run, in which case the new content is appended. I suggest to do the same for the HDF5 files here. The corresponding function to call is IO_TruncateOutputFiles, defined by CactusBase/IOUtil.
-erik
users@lists.einsteintoolkit.org