#2542: Support for creating and saving 1D histograms
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component: Carpet
Comment (by Roland Haas):
Instead of just offering a signle type of reduction \(historgram\) it might be nicer \(and not more work\) to allow users to supply a custom “kernel” that is being called for each grid point by CarpetReduce. Something like:
```
int callback(cGH* cctkGH, CCTK_INT idx, CCTK_REAL data, void *calldata)
```
or so which would let users implement their own histogram or other reductions, eg. a maximum location reduction.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2542/support-for-creat…
#2543: Consolidate data formats to simplify postprocessing
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component:
Comment (by Gabriele Bozzola):
Thanks Wolfgang for this clear and detailed wishlist. I would like to add another item and share a thought.
10\. Grid arrays should be clearly distinguishable from grid functions. At the moment, Carpet outputs grid arrays in the exact same way as grid function. CarpetIOASCII also attaches some coordinates \(that are completely meaningless\). This confuses post-processing tools.
\(Bonus. Thorns should be discouraged from implementing their custom writing routines \(e.g., `VolumeIntegrals`\), unless strictly needed. \)
In the spirit of consolidating the data formats, I argue that Einstein Toolkit should provide a clear and well-documented way to analyze simulations. At the moment, Einstein Toolkit supports multiple different types of output and users write their own homebrew scripts depending on the way they decide to run simulations. For example, people might decide to output 2D grid data in the objectively inferior ASCII format because it is much easier to plot with `gnuplot`. If we provide a clear way to post-process data, users won’t need to understand the \(not clearly documented\) intricacies of how data is output to start doing science, and the way the output is stored will become an implementation detail under our control. This will allow us to settle on a specification for how the output should be produced.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2543/consolidate-data-…
#2543: Consolidate data formats to simplify postprocessing
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component:
Comment (by Roland Haas):
Instead of just offering a signle type of reduction \(historgram\) it might be nicer \(and not more work\) to allow users to supply a custom “kernel” that is being called for each grid point by CarpetReduce. Something like:
```
int callback(cGH* cctkGH, CCTK_INT idx, CCTK_REAL data, void *calldata)
```
or so which would let users implement their own histogram or other reductions, eg. a maximum location reduction.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2543/consolidate-data-…
#2543: Consolidate data formats to simplify postprocessing
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component:
Currently, writing postprocessing tools for ET is unnecessarily difficult because required information needs to be collected from many locations, has to accommodate competing standards, and sometimes require guesses using heuristics. Below is a list of improvements from the postprocessing viewpoint, which is not complete and can be augmented over time.
1. Table of content for grid variable output. Each output folder should contain a machine-readable file that keeps track of all files containing grid data, with a list of variables and the available timesteps for each variable. Of course this should distinguish between 1D, 2D and 3D output. Currently, one has to open all files and parse the content for metadata, which can be very slow with HDF5. The issue is especially problematic when using one file per group.
2. The same for reduction output.
3. A machine readable file with all parameters and their values, including those not set in the parfile and set to default. The values should be values, postprocessing code should not have to emulate the handmade programming language parfiles have become. Each folder with a restart should have one such file in a standard location/name.
4. The reductions thorn should also output enough information to convert norm1/average into volume integrals, i.e. a scalar x such that `volume integral = x * average`
5. Unique extensions. There should be one and only one unique extension for each type of file, across all standard thorns that produce output. In particular, just adding ‘.h5’ is not enough. For example 3D data currently has extension `xyz.h5` or just `.h5` and multipole data can have extension `.h5` as well.
6. Simfactory should also provide machine-readable metadata about restarts and simulation folders. It should be possible to easily obtain a tree-like structure of the various restarts, complete with iteration ranges.
7. One standard format for timeseries. Currently reductions, 0D output, multipoles, and AHorizonFinder all have their own formats, multipoles has even two formats. Timeseries files should also contain metadata with the range of available iterations and times.
8. Settle on one format for each type of data and deprecate the rest, unless several formats are really needed, e.g. for performance reasons. A prime example is 2D ascii output, which is inferior to hdf5 in every way. If deprecation is not possible, having tools to convert from all competing formats into a canonical format after the simulation might help. Another duplication of effort is caused by the one-per-group/one-per-variable duality.
9. There should support to add arbitrary metadata to simulations. For example, initial data thorns could add model properties such as BNS masses, spins, separations. There should be an API such that code can add metadata, and some standard location where all metadata is collected in a machine and human readable format such as json. This should be on the simulation level, such that new metadata can be added during restarts, but existing metadata is immutable. Each thorn should have its own metadata namespace.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2543/consolidate-data-…
#2542: Support for creating and saving 1D histograms
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component: Carpet
One common task that is not easy in ET is to create simple one-dimensional histograms from grid variables. There should be support for this in the core infrastructure. Writing user thorns for this task requires the use of internal implementation details such as the Carpet weight mask. This also ties such code to Carpet and prevents use with e.g. CarpetX. Further, there needs to be a standard to save such histograms to disk at regular timesteps, much like the grid data itself.
Proposal:
1. There should be a histogram API that allows creating histograms, very much like 1D arrays. The IO thorns should write those to disk at user specified timesteps, similarly as done for arrays. In addition to the bins, there should be two special bins for values above and below the bin range. Metadata that needs to be saved as well is the binning \(lower bound, upper bound, number of bins\). Specifying histograms should be possible via parfile, populating them should be possible for user code via function calls.
2. There should be a thorn that allows to create arbitrarily many histograms, each from two grid variables, one used for binning, and one for the weight. The weight should be a density with respect to coordinate volume. The cell size and weights from refinement masks should be handled by the thorn, not the user.
Example applications:
A histogram of ejecta velocities in BNS merger simulations. This requires a one-dimensional histogram at a given timestep, where the bins correspond to velocity. The baryonic mass in each cell then needs to be added in the bin given by the velocity in the same cell. All the user would have to do is write a thorn that sets up two scalar grid variables, one with unbound mass density \(per coordinate volume\) and one with the velocity.
Another thorn might use a different approach for measuring the same, computing fluxes through a surface instead. Such a thorn could then still use the histogram API directly to get standardized file IO of those histograms.
A BNS tracking thorn may employ two histograms of masses binned by phi-position and radial coordinate, respectively, to determine orbital phase and separation without resorting to fragile code that resorts to reduction operations.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2542/support-for-creat…