Hello, I'm interested in ways to manipulate carpet hdf5 data from multiple jobs (e.g. from different output-XXXX directories of a simfactory-managed simulation) for visualization purposes.
Currently, if I had a variable I wanted to plot, what I would do is run hdf5_merge on the separate .h5 files to generate a single hdf5 file and then use a visualization tool (like visit) with it.
I'm curious if there is a means to do this differently: for instance, if one could load the separate .h5 files in python as some collection of data objects and then combine those objects, no separate and combined .h5 file would need to be produced. Then, if some means existed to create a visualization from that data object (e.g. through yt), the need for the combined .h5 file would be totally obsolete.
I suspect this must already be possible in some manner, perhaps using other programs. I'd be curious to hear how, for instance, SimulationTools handles this problem.
Hello Michael,
I wrote a small python library that does exactly what you are looking for. You can use it to read a set of HDF5 and it returns an object that can be queried for grid functions at particular times and/or locations. The code is publicly available here:
https://bitbucket.org/dradice/scidata
The repository contains a lot of python modules, most written by me early on during my PhD studies (that is to say most are crap), but the ones that might be interesting for you are:
scidata/carpet/hdf5.py scidata/carpet/grid.py
These have been rewritten many times and they are quite efficient at reading large datasets quickly. We have used them to analyze the simulations from (arXiv:1512.00838), which produced ~200 Tb of data scattered through hundreds of thousands h5 files.
You can see the examples/field.py for an example with 2D data, the handling of 1D and 3D data is identical. Once you have read the data, you can plot or analyze it with your favorite tools (e.g. matplotlib, or yt).
The main caveat is that the code is python-2.7, maybe python-2.6, only. I have not yet found the time / motivation to convert everything to python-3.
Best,
David
On May 17, 2016, at 9:52 AM, Michael Clark michael.clark@gatech.edu wrote:
Hello, I'm interested in ways to manipulate carpet hdf5 data from multiple jobs (e.g. from different output-XXXX directories of a simfactory-managed simulation) for visualization purposes.
Currently, if I had a variable I wanted to plot, what I would do is run hdf5_merge on the separate .h5 files to generate a single hdf5 file and then use a visualization tool (like visit) with it.
I'm curious if there is a means to do this differently: for instance, if one could load the separate .h5 files in python as some collection of data objects and then combine those objects, no separate and combined .h5 file would need to be produced. Then, if some means existed to create a visualization from that data object (e.g. through yt), the need for the combined .h5 file would be totally obsolete.
I suspect this must already be possible in some manner, perhaps using other programs. I'd be curious to hear how, for instance, SimulationTools handles this problem. _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On Tue, May 17, 2016 at 04:52:10PM +0000, Michael Clark wrote:
Hello, I'm interested in ways to manipulate carpet hdf5 data from multiple jobs (e.g. from different output-XXXX directories of a simfactory-managed simulation) for visualization purposes.
Currently, if I had a variable I wanted to plot, what I would do is run hdf5_merge on the separate .h5 files to generate a single hdf5 file and then use a visualization tool (like visit) with it.
I'm curious if there is a means to do this differently: for instance, if one could load the separate .h5 files in python as some collection of data objects and then combine those objects, no separate and combined .h5 file would need to be produced. Then, if some means existed to create a visualization from that data object (e.g. through yt), the need for the combined .h5 file would be totally obsolete.
I suspect this must already be possible in some manner, perhaps using other programs. I'd be curious to hear how, for instance, SimulationTools handles this problem.
Try the attached script. It uses python to get a list of datasets from however many hdf5 files you specify and creates a new (small) hdf5 file that contains only links to the original files. This can be used to, e.g., only specify a single file in VisIt, but can also be used (and is, if you look into the script), to 'filter' some datasets (timesteps, refinement levels, whatever you like) which VisIt wouldn't see, and thus, not initially load, reducing load times by a lot in some cases.
The primary use, however, and the reason I wrote this, was that this way you can combine variables that had been written to multiple files to one, allowing much easier manipulation in VisIt, e.g., for vector components like the hydro velocity. It essentially provides a different 'view' of the same simulation to visit.
Note that because of that filtering, you might want to have a look at the (very short) script and adapt it to your needs before you use it.
Also, in case you have an old version of VisIt, you might need to enable the use of soft links in the VisIt Carpet reader. I am currently not sure if that change made it back upstream yet. The corresponding patch is:
Index: avtCarpetHDF5FileFormat.C =================================================================== --- avtCarpetHDF5FileFormat.C (revision 57) +++ avtCarpetHDF5FileFormat.C (working copy) @@ -1044,7 +1044,7 @@ sprintf(fullname, "%s%s%s", rootname, rootname[strlen(rootname)-1]=='/' ? "" : "/",member_name);
// we are interested in datasets only - skip anything else - H5Gget_objinfo (group_id, member_name, 0, &object_info); + H5Gget_objinfo (group_id, member_name, 1, &object_info); if (object_info.type != H5G_DATASET) { if (object_info.type == H5G_GROUP)
Frank
On Tue, May 17, 2016 at 12:19:09PM -0500, Frank Loeffler wrote:
Try the attached script. It uses python to get a list of datasets from however many hdf5 files you specify and creates a new (small) hdf5 file that contains only links to the original files. This can be used to, e.g., only specify a single file in VisIt
For some reason, I was thinking of VisIt when I read your email. If you want to use python to process your data anyway, then an intermediate file like this doesn't quite make sense. I'd suggest something like what David suggested. The Parma group has a similar python framework, and if I remember correctly, Wolfgang Kastaun as well.
The one of the Parma group can be found here:
https://einstein.pr.infn.it/svn/numrel/pub/PyCactus/ (anonymous / anon)
Among other things, it 'recombines' checkpoints, contains some post-processing routines, e.g., for gw extraction, and so on. It also reads (some of the) ASCII format, besides the obvious hdf5. It is intended to be used as utility-library for other, more specific post-processing scripts, e.g., plots for papers ect.
Frank
Hello Frank, Michael,
@Michael: both of what Frank or David describe should work. If you know how to compile the VisIt reader yourselv, you can also try the copy of CarpetHDF5 at https://bitbucket.org/rhaas80/carpethdf5 either the "multifile" branch (which is the default one that has all features) or the "for_VisIt" branch which has fewer features and is what I wanted at one point propose for inclusion in VisIt. Both will let you create .visit files listing your hdf5 files and VisIt will treat them a a single database. You will have to create .visit files that look like this:
!NBLOCKS 3 output-0000/rho.file_0.h5 output-0000/rho.file_1.h5 output-0000/rho.file_2.h5 output-0001/rho.file_0.h5 output-0001/rho.file_1.h5 output-0001/rho.file_2.h5
for a 3 process run and similar for runs with more processes. You can probably (I have not tested this) also list both rho and eps:
!NBLOCKS 3 output-0000/rho.file_0.h5 output-0000/rho.file_1.h5 output-0000/rho.file_2.h5 output-0000/eps.file_0.h5 output-0000/eps.file_1.h5 output-0000/eps.file_2.h5 output-0001/rho.file_0.h5 output-0001/rho.file_1.h5 output-0001/rho.file_2.h5 output-0001/eps.file_0.h5 output-0001/eps.file_1.h5 output-0001/eps.file_2.h5
possibly you have to set NBLOCKS to 6 (2*3) for this, though I have not tested this either.
@Frank:
Also, in case you have an old version of VisIt, you might need to enable the use of soft links in the VisIt Carpet reader. I am currently not sure if that change made it back upstream yet. The corresponding patch is:
Index: avtCarpetHDF5FileFormat.C
--- avtCarpetHDF5FileFormat.C (revision 57) +++ avtCarpetHDF5FileFormat.C (working copy) @@ -1044,7 +1044,7 @@ sprintf(fullname, "%s%s%s", rootname, rootname[strlen(rootname)-1]=='/' ? "" : "/",member_name);
// we are interested in datasets only - skip anything else
- H5Gget_objinfo (group_id, member_name, 0, &object_info);
- H5Gget_objinfo (group_id, member_name, 1, &object_info); if (object_info.type != H5G_DATASET) { if (object_info.type == H5G_GROUP)
As far as I can tell this code change has not even made it into the ET maintained copy of the reader: https://svn.cactuscode.org/VizTools/CarpetHDF5/
where if this does not cause any other problems I think it would be welcome as well.
Yours, Roland
Hello,
I'd like to point out also my python library/script called "rugutils". This one started off inspired by David's scidata (and pygraph), but has developed into a set of tools to look in a quick and hassle-free way at Carpet data. It's written in python 3.5
https://bitbucket.org/fguercilena/rugutils https://docs.einsteintoolkit.org/et-docs/Analysis_and_post-processing
Federico
2016-05-17 19:44 GMT+02:00 Roland Haas rhaas@aei.mpg.de:
Hello Frank, Michael,
@Michael: both of what Frank or David describe should work. If you know how to compile the VisIt reader yourselv, you can also try the copy of CarpetHDF5 at https://bitbucket.org/rhaas80/carpethdf5 either the "multifile" branch (which is the default one that has all features) or the "for_VisIt" branch which has fewer features and is what I wanted at one point propose for inclusion in VisIt. Both will let you create .visit files listing your hdf5 files and VisIt will treat them a a single database. You will have to create .visit files that look like this:
!NBLOCKS 3 output-0000/rho.file_0.h5 output-0000/rho.file_1.h5 output-0000/rho.file_2.h5 output-0001/rho.file_0.h5 output-0001/rho.file_1.h5 output-0001/rho.file_2.h5
for a 3 process run and similar for runs with more processes. You can probably (I have not tested this) also list both rho and eps:
!NBLOCKS 3 output-0000/rho.file_0.h5 output-0000/rho.file_1.h5 output-0000/rho.file_2.h5 output-0000/eps.file_0.h5 output-0000/eps.file_1.h5 output-0000/eps.file_2.h5 output-0001/rho.file_0.h5 output-0001/rho.file_1.h5 output-0001/rho.file_2.h5 output-0001/eps.file_0.h5 output-0001/eps.file_1.h5 output-0001/eps.file_2.h5
possibly you have to set NBLOCKS to 6 (2*3) for this, though I have not tested this either.
@Frank:
Also, in case you have an old version of VisIt, you might need to enable the use of soft links in the VisIt Carpet reader. I am currently not sure if that change made it back upstream yet. The corresponding patch is:
Index: avtCarpetHDF5FileFormat.C
--- avtCarpetHDF5FileFormat.C (revision 57) +++ avtCarpetHDF5FileFormat.C (working copy) @@ -1044,7 +1044,7 @@ sprintf(fullname, "%s%s%s", rootname,
rootname[strlen(rootname)-1]=='/' ? "" : "/",member_name);
// we are interested in datasets only - skip anything else
- H5Gget_objinfo (group_id, member_name, 0, &object_info);
- H5Gget_objinfo (group_id, member_name, 1, &object_info); if (object_info.type != H5G_DATASET) { if (object_info.type == H5G_GROUP)
As far as I can tell this code change has not even made it into the ET maintained copy of the reader: https://svn.cactuscode.org/VizTools/CarpetHDF5/
where if this does not cause any other problems I think it would be welcome as well.
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
It is surprising to see how many such frameworks are available. I'd suggest that we create a wiki page that briefly describes them, and points to instructions for downloading and examples.
Incidentally, we are also developing such a framework here at Perimeter. The basic idea is to have a generic way of describing the data in a simulation, introducing concepts such as discretizations of a manifold, bases for the tangent space, tensor types, fields, etc. < https://github.com/eschnett/SimulationIO%3E. We (i.e. Jonah Miller) are now testing and benchmarking it against real-world data, developing a yt < http://yt-project.org%3E reader in the process.
The SimulationIO library is supposed to be used directly for writing and reading data from Cactus. Currently, that hasn't been implemented yet, and so there is a converter from Carpet output that either collects all the metadata (with external links to the original datasets) or creates a single output file. Obviously the file format is still based on HDF5 -- the main advantage is that the metadata are much easier to access and interpret than having to loop over the components and parsing their names.
I assume Jonah will describe this at the Einstein Toolkit.
-erik
On Tue, May 17, 2016 at 2:17 PM, Federico Guercilena < guercilena@th.physik.uni-frankfurt.de> wrote:
Hello,
I'd like to point out also my python library/script called "rugutils". This one started off inspired by David's scidata (and pygraph), but has developed into a set of tools to look in a quick and hassle-free way at Carpet data. It's written in python 3.5
https://bitbucket.org/fguercilena/rugutils https://docs.einsteintoolkit.org/et-docs/Analysis_and_post-processing
Federico
2016-05-17 19:44 GMT+02:00 Roland Haas rhaas@aei.mpg.de:
Hello Frank, Michael,
@Michael: both of what Frank or David describe should work. If you know how to compile the VisIt reader yourselv, you can also try the copy of CarpetHDF5 at https://bitbucket.org/rhaas80/carpethdf5 either the "multifile" branch (which is the default one that has all features) or the "for_VisIt" branch which has fewer features and is what I wanted at one point propose for inclusion in VisIt. Both will let you create .visit files listing your hdf5 files and VisIt will treat them a a single database. You will have to create .visit files that look like this:
!NBLOCKS 3 output-0000/rho.file_0.h5 output-0000/rho.file_1.h5 output-0000/rho.file_2.h5 output-0001/rho.file_0.h5 output-0001/rho.file_1.h5 output-0001/rho.file_2.h5
for a 3 process run and similar for runs with more processes. You can probably (I have not tested this) also list both rho and eps:
!NBLOCKS 3 output-0000/rho.file_0.h5 output-0000/rho.file_1.h5 output-0000/rho.file_2.h5 output-0000/eps.file_0.h5 output-0000/eps.file_1.h5 output-0000/eps.file_2.h5 output-0001/rho.file_0.h5 output-0001/rho.file_1.h5 output-0001/rho.file_2.h5 output-0001/eps.file_0.h5 output-0001/eps.file_1.h5 output-0001/eps.file_2.h5
possibly you have to set NBLOCKS to 6 (2*3) for this, though I have not tested this either.
@Frank:
Also, in case you have an old version of VisIt, you might need to enable the use of soft links in the VisIt Carpet reader. I am currently not sure if that change made it back upstream yet. The corresponding patch is:
Index: avtCarpetHDF5FileFormat.C
--- avtCarpetHDF5FileFormat.C (revision 57) +++ avtCarpetHDF5FileFormat.C (working copy) @@ -1044,7 +1044,7 @@ sprintf(fullname, "%s%s%s", rootname,
rootname[strlen(rootname)-1]=='/' ? "" : "/",member_name);
// we are interested in datasets only - skip anything else
- H5Gget_objinfo (group_id, member_name, 0, &object_info);
- H5Gget_objinfo (group_id, member_name, 1, &object_info); if (object_info.type != H5G_DATASET) { if (object_info.type == H5G_GROUP)
As far as I can tell this code change has not even made it into the ET maintained copy of the reader: https://svn.cactuscode.org/VizTools/CarpetHDF5/
where if this does not cause any other problems I think it would be welcome as well.
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Federico Guercilena Institut für Theoretische Physik Johann Wolfgang Goethe-Universität Max-von-Laue-Str. 1 60438 Frankfurt am Main, Germany Telephone: +49 69 798 47887 Email: guercilena[at]th.physik.uni-frankfurt.de
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Hello all,
It is surprising to see how many such frameworks are available. I'd suggest that we create a wiki page that briefly describes them, and points to instructions for downloading and examples.
Such a page exists (I had the same thought as Erik and noticed the page, from its creation date it stems from last year's ET workshop in Stockholm, kudos to Ian Hinder who created it): https://docs.einsteintoolkit.org/et-docs/Analysis_and_post-processing
Yours, Roland
On Tue, May 17, 2016 at 03:08:18PM -0400, Erik Schnetter wrote:
It is surprising to see how many such frameworks are available. I'd suggest that we create a wiki page that briefly describes them, and points to instructions for downloading and examples.
We already have that: https://docs.einsteintoolkit.org/et-docs/Analysis_and_post-processing
Frank
On Tue, May 17, 2016 at 12:52 PM, Michael Clark michael.clark@gatech.edu wrote:
I'm curious if there is a means to do this differently: for instance, if one could load the separate .h5 files in python as some collection of data objects and then combine those objects, no separate and combined .h5 file would need to be produced. Then, if some means existed to create a visualization from that data object (e.g. through yt), the need for the combined .h5 file would be totally obsolete.
I suspect this must already be possible in some manner, perhaps using other programs. I'd be curious to hear how, for instance, SimulationTools handles this problem.
This is exactly what SimulationTools does. It abstracts away the fact that there are potentially many datasets scattered across many HDF5 files (e.g. *.file_N.h5) scattered across multiple restarts. You just ask it to read a specified grid function from a simulation at a specified iteration and it works out the details. It achieves this by scanning the list of datasets across all files in a simulation and determining which datasets must be loaded from which files for a given grid function, iteration, refinement level, etc. It then loads the data from those datasets and merges the components together into a single object (which we call a DataRegion) that represents the data for that grid function at that iteration.
Barry
users@lists.einsteintoolkit.org