#2546: CarpetIOHDF5: Don't set "delta" attribute to zero for zero-width components
Reporter: Erik Schnetter
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component:
Comment (by Roland Haas):
Please review.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2546/carpetiohdf5-dont…
#2546: CarpetIOHDF5: Don't set "delta" attribute to zero for zero-width components
Reporter: Erik Schnetter
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component:
Carpet’s HDF5 format outputs a `delta` attribute that describes the grid spacing for each written chunk. Currently `delta` is always set to zero for zero-width components, i.e. e.g. for 2d arrays. This confuses post-processing tools.
This pull request passes the actual grid spacing to the writer routine, and writes the correct grid spacing into the attribute.
The pull request is at [https://bitbucket.org/eschnett/carpet/pull-requests/46](https://bitbucket.o… .
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2546/carpetiohdf5-dont…
#2543: Consolidate data formats to simplify postprocessing
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component:
Comment (by Erik Schnetter):
Wolfgang: I see I misread you question about 0D data and reductions. Yes, there is one file per iteration with all reductions. This file is small, and it’s probably more efficient to have them in a single file.
0D output isn’t implemented yet. 1D output is currently \(very inefficiently\) one file per iteration per variable per direction. My plan is to have 1D and 2D output in binary as well since it’s not scalable to keep it in ASCII. ASCII output is very convenient for debugging, I don’t have a good solution for this yet. Maybe keep the current inefficient format around, or have a simple tool that converts binary to ASCII.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2543/consolidate-data-…
#2543: Consolidate data formats to simplify postprocessing
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component:
Comment (by Erik Schnetter):
The idea is to have one such metafile per output directory. I’ll have to think about the one-file-per-iteration setup, it does seem wasteful and inconvenient. There certainly won’t be one file per node, things would be aggregated.
If you like the design for CarpetX, then we can repeat it for Carpet \(or, rather, CactusBase/IOUtil\) and use it for all simulations. I’m sure we’ll have to iterate on the design, and then flush out a few bugs where the design doesn’t make sense or is too limited.
0D output etc. are already supported in the format. There is a key that specifies which directions of a variable is output: `[0,1,2]` is for 3D output, `[1]` is for output in the y direction, etc.
Yes, HDF5 is slow. That’s why I want to switch to ADIOS2. This format has essentially the same capabilities as HDF5 \(blocks of data, arbitrary types, attributes, groups, etc.\), but it properly separates metadata, and is parallel by design. It’s also safe to append to an ADIOS2 file \(internally data are separated into “iterations” which cannot be modified once written\).
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2543/consolidate-data-…
#2543: Consolidate data formats to simplify postprocessing
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component:
Comment (by Wolfgang Kastaun):
Regarding the new format for reductions, do I understand correctly that there will be one file per iteration with all available reductions? That means one has to parse the data for all reductions just to get one of them. However, given the small data size, this might not be too wasteful.
What about other data of type time series, e.g. 0D output, are they treated different, format-wise?
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2543/consolidate-data-…
#2543: Consolidate data formats to simplify postprocessing
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component:
Comment (by Wolfgang Kastaun):
The proposed solution for CarpetX would probably simplify the logic which gathers information on the available data a lot. I’m not entirely sure about speed. The reason parsing the hdf5 files takes so long seems a design flaw that requires basically to read the whole file just to get the names of all datasets, combined with the unfortunate choice of having the table of content information only in those names. I also did not get how the solution would look like for runs using many nodes, would there be one file for each node or even MPI process, or is the information collected first on one node?
But in principle, reading a few thousand yaml files should not take that long, we would have to try.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2543/consolidate-data-…
#2543: Consolidate data formats to simplify postprocessing
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component:
Comment (by Wolfgang Kastaun):
Maybe we should have two tickets, one for CarpetX and one for simplifying postprocessing with the Carpet-based infrastructure, assuming that will stay around for some time. It will probably take a while until I update postcactus for CarpetX.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2543/consolidate-data-…