#2543: Consolidate data formats to simplify postprocessing
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component:
Comment (by Erik Schnetter):
The key `format_name` and `format_version` tell you how to interpret the content of the file. I don’t think it is feasible to have a generic description in the metadata that would allow you to extract the information from all the file formats we support. \(If that was possible then we wouldn’t have multiple file formats.\) Thus the reader needs to understand and have special support for the specific file format, be it Silo or HDF5 or TSV or JSON.
In this particular case, the TSV file has column headers that identify the content. The TSV files are small, and reading them is fast.
Regarding slow reading: In the example I provided, there are two files written per iteration, one TSV file with the norms, and one Silo/HDF5 file with 3D data. If there had been many nodes, then there might have been more files to allow nodes to write data independently \(which speeds things up\).
If there is a thorn that produces many files, then the thorn should be updated, or it should switch to a standard output format. That issue is independent of setting up a table of contents. There are probably cases where having many ASCII files makes sense, e.g. for debugging or developing scripts.
I chose yaml since that is already supported in CarpetX, and because it is human-readable. The file encoding really isn’t important and can be changed quickly.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2543/consolidate-data-…
#2543: Consolidate data formats to simplify postprocessing
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component:
Comment (by Gabriele Bozzola):
I can start giving you some first comments, but one would have to start thinking about the design of the postprocessing tool to make more serious comments.
One quick comment is that the file doesn’t tell me all how to read a variable. Suppose I want to read `max(hydrogpu::rho)`, what column is it in the tsv file? If I understand how the tsv is structured \(it contains many reductions\), we need to enforce certain constrains to make sure that we can determine the column number without parsing the tsv file. For example, the order of variables and reductions in the yaml file must be the same as in the tsv file and all the variables must appear in all the reductions.
A second comment is: on some shared filesystems, opening files can be extremely expensive, so having fewer big files is much better than having many small ones. An example of this is the how Einstein Toolkit outputs multipoles now, which can lead to thousands of ASCII files, or one HDF5 file \(and the performance difference is really important\). If each the data for each iteration is stored in different files, I worry that it might lead to performance problems.
Also, why did you pick yaml over the \(faster but less powerful\) json?
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2543/consolidate-data-…
#2543: Consolidate data formats to simplify postprocessing
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component:
Comment (by Erik Schnetter):
I have added a routine to CarpetX that output metadata similar to the ones you request. See [here](https://gist.github.com/eschnett/8ac85b09f0c156fe7c224d8420aa37d4) for an example.
It is difficult to create a single file that contains information about all iterations, because this means that this file needs to be overwritten at each iteration. Instead, I’m creating a new file per iteration, and the post-processing tool should read all these files.
Since this is just a proof of concept, there are two metadata files \(that should change later\). The first, `cactus-metadata`, should have all the data requested above. The second, `carpetx-metadata`, already existed in CarpetX. It describes the grid structure \(that’s probably not interesting here\), but incidentally also all parameters in the very explidit format requested.
The metadata files are output in yaml. That’s a standard format that should be easy to read in Python, resulting in dictionaries and arrays.
Wolfgang, Gabriele, could you comment?
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2543/consolidate-data-…
#2544: utf-8 decode failure when using create-run
Reporter:
Status: open
Milestone:
Version: development version
Type: bug
Priority: major
Component: SimFactory
Changes (by Roland Haas):
assignee: Bruno Giacomazzo (was )
responsible: [] (was )
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2544/utf-8-decode-fail…
#2545: some of simfactory's run scripts use /bin/sh when they really want /bin/bash
Reporter:
Status: new
Milestone:
Version: development version
Type: bug
Priority: minor
Component: SimFactory
On some systems /bin/sh is not bash but instead "dash" which is more lightweight but still a POSIX shell. Simfactory's run scripts usually are written with bash in mind so should use /bin/bash and not /bin/sh
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2545/some-of-simfactor…
#2544: utf-8 decode failure when using create-run
Reporter:
Status: new
Milestone:
Version: development version
Type: bug
Priority: major
Component: SimFactory
Lorenzo Ennoggi just encountered an issue with simfactory where it fails with an utf-8 decoding error due to an invalid utf-8 character in Cactus' output stream in simfactory/lib/simrestart.py line 1016:
```
out_txt = out_read.read()
```
then can probably be avoided by ensuring that we read and write "byte" objects (and use binary file IO) rather than "strings".
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2544/utf-8-decode-fail…
#2542: Support for creating and saving 1D histograms
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component: Carpet
Comment (by Wolfgang Kastaun):
Such a low level kernel may be a good building block for implementing the histograms, but is only addressing part of the problem. The callback function could store histograms of local data, but this still has to be reduced globally at some point. When exactly to schedule the reduction and the resetting of the histograms between timesteps will then probably become another dark art \(scheduling GLOBAL-LATE or GLOBAL-EARLY and what bin?\). It also does not handle output in a standard format, nor the basic histogram code. In short, I think histograms are such a basic feature that they should be very easy to create for the user, similar to scalar reductions.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2542/support-for-creat…