Hi,
We had quite a discussion about the problems the current visualization tools have reading our hdf5 format. It became clear that we have to change/extend it, but before we think about how, I would like to see what specifically the readers have problems with. Let me start here with what I can recall, and please freely add to that.
One of the problems is that readers currently have to iterate over all datasets present to get to some information they need. Getting the list of all datasets by itself can take quite a while on a regular hdf5 file, but then readers also have to look at attributed within each dataset. While of of this is necessary to visualize all data in a given file, most of the time not all of the data is actually necessary, and certainly not for just 'opening the file'.
operations a reader needs to be fast: - list of variables (at a given iteration) - list of time steps / iterations - AMR structure for one given iteration (all maps, rls and components)
Regardless of these additional meta-data, we already established the need for a meta-data file, effectively a copy of all datasets but without the actual data. Am I remembering correctly that the idea was to write this at run-time (eventually - right now it could be generated as post-processing)?
Is this all readers would need?
Frank
Hi Frank,
Erik can probably elaborate more, but my experience with yt has been that it is very difficult to correctly reconstruct the Carpet grid structure from our hdf5 output. Therefore, perhaps it would be nice to have more substantive (and easy to parse) grid information output at each time step.
Perhaps ask the developers of various post-processing and visualization toolboxes and ask them what information makes reading simpler.
Best, Jonah Miller
On 15-08-18 04:02 PM, Frank Loeffler wrote:
Hi,
We had quite a discussion about the problems the current visualization tools have reading our hdf5 format. It became clear that we have to change/extend it, but before we think about how, I would like to see what specifically the readers have problems with. Let me start here with what I can recall, and please freely add to that.
One of the problems is that readers currently have to iterate over all datasets present to get to some information they need. Getting the list of all datasets by itself can take quite a while on a regular hdf5 file, but then readers also have to look at attributed within each dataset. While of of this is necessary to visualize all data in a given file, most of the time not all of the data is actually necessary, and certainly not for just 'opening the file'.
operations a reader needs to be fast:
- list of variables (at a given iteration)
- list of time steps / iterations
- AMR structure for one given iteration (all maps, rls and components)
Regardless of these additional meta-data, we already established the need for a meta-data file, effectively a copy of all datasets but without the actual data. Am I remembering correctly that the idea was to write this at run-time (eventually - right now it could be generated as post-processing)?
Is this all readers would need?
Frank
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On Tue, Aug 18, 2015 at 4:02 PM, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi,
We had quite a discussion about the problems the current visualization tools have reading our hdf5 format. It became clear that we have to change/extend it, but before we think about how, I would like to see what specifically the readers have problems with. Let me start here with what I can recall, and please freely add to that.
One of the problems is that readers currently have to iterate over all datasets present to get to some information they need. Getting the list of all datasets by itself can take quite a while on a regular hdf5 file, but then readers also have to look at attributed within each dataset. While of of this is necessary to visualize all data in a given file, most of the time not all of the data is actually necessary, and certainly not for just 'opening the file'.
operations a reader needs to be fast:
- list of variables (at a given iteration)
- list of time steps / iterations
- AMR structure for one given iteration (all maps, rls and components)
Regardless of these additional meta-data, we already established the need for a meta-data file, effectively a copy of all datasets but without the actual data. Am I remembering correctly that the idea was to write this at run-time (eventually - right now it could be generated as post-processing)?
Is this all readers would need?
Frank
As Jonah described during the ET workshop, we're working with the yt developers to make Carpet's HDF5 format yt-friendly. At the moment, we are adding the missing information, which so far included the set of active grid points -- i.e. those that should be displayed for a given level, as opposed to those that should be cut off (ghost, buffer, symmetry points). One may argue that these points should not have been output at all, but that would be a major change to the current file format.
Additional information may also be useful. This is information that is already present in some way, but is cumbersome to extract. Presumably, Carpet already reconstructs this information when reading data from a file, but a visualization or analysis tools may want to access information in a different order, and may not want to reimplement most of CarpetIOHDF5.
The items you list are a good starting point, but are too high-level to be useful as a guide to implementing this. To make things concrete, I'd rather collaborate with someone who is actually implementing a reader, and provide the data that this reader needs. For example, you say you want a "list of variables (as a given iteration)" -- do you really want a two-step table, containing first a list of iterations, and then (for each iteration) a list of variables? That's likely very different from what a user wants to extract; rather, people want a set of variables, and for each variable, the set of iterations at which this variable has data. Given that we have an AMR format where certain levels exist only at certain iterations, this requires a bit more detail for a complete specification.
How do you want the AMR structure to be presented? Currently, Carpet can output a string that can be parsed reasonably easily, and which describes the grid structure. Is that sufficient?
I hear the visualization tools are interested in learning the connectivity between the different refinement levels. How should this be represented? With vertex centering, the vertices overlap, and if you want to display the cells around each vertex (control volume, necessary for volume rendering), then you need to handle partial volumes. With cell centering, things are easier, at least if you use CarpetRegrid2::snap_to_coarse -- which Carpet doesn't require, but the visualization tool may want to insist on it.
So, in addition to giving a high-level list of features, let's also collect a list of tools that will help us find out the details about these wanted features.
We also may want to have a mechanism to "glue" the different output file from different processors together, other than just looking for files with similar names.
Finally, I disagree with the "established the need for a meta-data file". This capability exists, and it speeds up reading for the current output format, but the current output format has several obvious shortcomings; if those were remedied, things may be much faster.
-erik
Hi,
I am not a visualization expert, so I am only answering what I know. Please others: chime in. I only wanted to get the discussion going and avoid different people working on the same thing separately after the workshop.
On Wed, Aug 19, 2015 at 11:51:00AM -0400, Erik Schnetter wrote:
As Jonah described during the ET workshop, we're working with the yt developers to make Carpet's HDF5 format yt-friendly. At the moment, we are adding the missing information,
Great. It might be good to share this with others.
which so far included the set of active grid points -- i.e. those that should be displayed for a given level, as opposed to those that should be cut off (ghost, buffer, symmetry points). One may argue that these points should not have been output at all, but that would be a major change to the current file format.
We already have parameters to do that.
The items you list are a good starting point, but are too high-level to be useful as a guide to implementing this. To make things concrete, I'd rather collaborate with someone who is actually implementing a reader, and provide the data that this reader needs. For example, you say you want a "list of variables (as a given iteration)" -- do you really want a two-step table, containing first a list of iterations, and then (for each iteration) a list of variables?
That is the kind of discussion I wanted to get started.
That's likely very different from what a user wants to extract; rather, people want a set of variables, and for each variable, the set of iterations at which this variable has data.
I agree.
How do you want the AMR structure to be presented? Currently, Carpet can output a string that can be parsed reasonably easily, and which describes the grid structure. Is that sufficient?
I didn't try this myself, but I remember someone at the workshop mentioning parsing the string as 'quite complicated' (no quote). Does this string also include information about all the components, so that a reader can easily figure out which component is containing a specific point/region?
We also may want to have a mechanism to "glue" the different output file from different processors together, other than just looking for files with similar names.
I experimented with external links in hdf5. They work rather nicely (after changing one byte in the source of the VisIt reader: enabling following links). I use this to have one (very small) hdf5 file pointing to all three components of a vector, so that it is easy to combine them to a vector in VisIt (since then from the point of visit they come from the same database). The 'recombiner' for this is a very tiny python script, but this could be done during a simulation as well.
Finally, I disagree with the "established the need for a meta-data file". This capability exists, and it speeds up reading for the current output format, but the current output format has several obvious shortcomings; if those were remedied, things may be much faster.
The main idea this reasoning comes from is that if meta-data and data are written intermixed, then just reading the meta-data is always going to be slower than if it would be contained in a file missing the actual data, due to disk access usually reading much more than requested. This is separate from what the meta data actually looks like.
Frank
On Wed, Aug 19, 2015 at 12:46 PM, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi,
I am not a visualization expert, so I am only answering what I know. Please others: chime in. I only wanted to get the discussion going and avoid different people working on the same thing separately after the workshop.
On Wed, Aug 19, 2015 at 11:51:00AM -0400, Erik Schnetter wrote:
As Jonah described during the ET workshop, we're working with the yt developers to make Carpet's HDF5 format yt-friendly. At the moment, we
are
adding the missing information,
Great. It might be good to share this with others.
which so far included the set of active grid points -- i.e. those that should be displayed for a given level, as opposed to those that should be cut off (ghost, buffer, symmetry points). One may argue that these points should not have been output at all, but that would be a major change to the current file format.
We already have parameters to do that.
The items you list are a good starting point, but are too high-level to
be
useful as a guide to implementing this. To make things concrete, I'd
rather
collaborate with someone who is actually implementing a reader, and
provide
the data that this reader needs. For example, you say you want a "list of variables (as a given iteration)" -- do you really want a two-step table, containing first a list of iterations, and then (for each iteration) a
list
of variables?
That is the kind of discussion I wanted to get started.
That's likely very different from what a user wants to extract; rather, people want a set of variables, and for each variable,
the
set of iterations at which this variable has data.
I agree.
How do you want the AMR structure to be presented? Currently, Carpet can output a string that can be parsed reasonably easily, and which describes the grid structure. Is that sufficient?
I didn't try this myself, but I remember someone at the workshop mentioning parsing the string as 'quite complicated' (no quote). Does this string also include information about all the components, so that a reader can easily figure out which component is containing a specific point/region?
The yt developers -- who probably have more of a computer science background -- were not bothered by the format.
Feel free to suggest a different format, or a different mechanism. The grid structure consists essentially of two nested arrays: A set of refinement levels, and each has a set of components. A component is described by the location of the origin and its size, i.e. two integer three-vectors. In C++, you would have something like
struct bbox { int origin[3], size[3]; }; typedef vector<vector<bbox>> grid_structure;
There's a slight complication if you have a multi-block system (adding either another level to the arrays, or adding an integer to the bbox structure describing the patch number). You also need to be careful about how the integer coordinates between the refinement levels are related, which is different for vertex and cell centred grids.
One way to store this without using a hierarchy would be to define instead
struct bbox { int reflevel; int patch; int origin[3]; int size[3]; };
and then store all such bboxes in a single, large array. It is the up to the reader to split this into a tree structure according to the reflevel and patch entries. That's easier to read, but more difficult to process.
-erik
We also may want to have a mechanism to "glue" the different output file
from different processors together, other than just looking for files
with
similar names.
I experimented with external links in hdf5. They work rather nicely (after changing one byte in the source of the VisIt reader: enabling following links). I use this to have one (very small) hdf5 file pointing to all three components of a vector, so that it is easy to combine them to a vector in VisIt (since then from the point of visit they come from the same database). The 'recombiner' for this is a very tiny python script, but this could be done during a simulation as well.
Finally, I disagree with the "established the need for a meta-data file". This capability exists, and it speeds up reading for the current output format, but the current output format has several obvious shortcomings;
if
those were remedied, things may be much faster.
The main idea this reasoning comes from is that if meta-data and data are written intermixed, then just reading the meta-data is always going to be slower than if it would be contained in a file missing the actual data, due to disk access usually reading much more than requested. This is separate from what the meta data actually looks like.
Frank
Hello all,
One of the problems is that readers currently have to iterate over all datasets present to get to some information they need. Getting the list of all datasets by itself can take quite a while on a regular hdf5 file, but then readers also have to look at attributed within each dataset. While of of this is necessary to visualize all data in a given file, most of the time not all of the data is actually necessary, and certainly not for just 'opening the file'.
For postprocessing needs what you usually need is a list of timesteps and in the file along with which variables and times exist in/correspond to those timesteps. For each datasset you then need to know pretty much all of the attributes to be able to properly place the dataset when reading it from disk and to decide which map/refinement level etc it is on. For postprocessing a simple table just just concatenates all attribute information from all datasets in all files that make out the output (eg for one set of file for one variable) is sufficient since one is usually fine even if this file becomes large (GB) since it is still small compared to the full dataset. For visualization/quick tests the GB sized file may be unsuitable.
A good compromise would be a (it seems to me) a hierarchical structure that is first indexed by timestep along with a table for times as a function of timestep (since the timestep is arbitrary and just a counter in Cactus). This would implicitly assume that the grid structure is identical between variables (which need not be true and in fact I had a case where this was not true). This works well for postprocessing since one typically steps through the timeseries sequentially.
An alternative that may be easier for visualization is to index by variable name first.
I agree that rather than make up some set of data out of thin air we should try and get some input from the target audience (visualization tools, postprocessing needs). For visualization I guess asking the yt developers seems like a good idea, we can also check what eg the VisIt reader needs to provide to VisIt (which is per timestep and per variable information I think) and we could try and see what the F5 format stores given that the author is also involved in visualization toolkits.
operations a reader needs to be fast:
- list of variables (at a given iteration)
- list of time steps / iterations
- AMR structure for one given iteration (all maps, rls and components)
All agreed.
Regardless of these additional meta-data, we already established the need for a meta-data file, effectively a copy of all datasets but
I would actually prefer if this information was included in the HDF5 files itself (replicated in each of them if possible) and written at the time the HDF5 file was created. We have some type of metadata in the index files that CarpetIOHDF5 can write yet they are not sufficient for very large datasets (since they do no reduce the number of files one has to parse).
without the actual data. Am I remembering correctly that the idea was to write this at run-time (eventually - right now it could be generated as post-processing)?
Yes, I would very much like if it was written at runtime.
Is this all readers would need?
I suspect that there will always be use cases where the information is arranged in just the "wrong" way to be convenient. So the best we can try is to support "typical" "well known" usage cases.
Yours, Roland
Hi
On Thu, Aug 20, 2015 at 09:49:52AM +0200, Roland Haas wrote:
A good compromise would be a (it seems to me) a hierarchical structure that is first indexed by timestep along with a table for times as a function of timestep (since the timestep is arbitrary and just a counter in Cactus). This would implicitly assume that the grid structure is identical between variables
Why would this assume that? In file system terms the first level of directories would just sort the components into their respective times.
(which need not be true and in fact I had a case where this was not true).
You have an example where the grid structure is different? Is this because you can specify, say, the refinement levels separately in the par file for each variable? That's an interesting use-case we should consider.
An alternative that may be easier for visualization is to index by variable name first.
We can have both. My current idea is to keep the current structure for backward compatibility, only adding meta data and structure (trees) which use symlinks to point to the actual components on the root level. Older readers could still read new files, newer readers could leverage the new information to be faster, and old files could even be 'upgraded' by adding the necessary new information.
I would actually prefer if this information was included in the HDF5 files itself (replicated in each of them if possible) and written at the time the HDF5 file was created. We have some type of metadata in the index files that CarpetIOHDF5 can write yet they are not sufficient for very large datasets (since they do no reduce the number of files one has to parse).
I see the value of that. In order to have this working efficiently, we could work with pre-allocating larger chunks for meta-data, and filling them bit by bit. This should (hopefully) make it faster than reading meta-data all over the output file.
Frank
users@lists.einsteintoolkit.org