#30: Implement a "get" command
--------------------------+-------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Resolution: | Keywords:
--------------------------+-------------------------------------------------
Comment (by knarf):
I might be interested in such a 'get' command, but personally I am not so
much concerned about consistency. As long as a simulation is running I
accept that files might be inconsistent between each other, and in the
case of hdf5 might also be broken at times. Pausing a simulation and
wasting SUs that way just to get a snapshot of the data 100% reliably
doesn't sound like a good idea to me. Automatically checking that an hdf5
file isn't broken would be nice though. However, making sure the file is
'current' seems unnecessary. I usually don't care to have the absolutely
last timestep if that was just written and would like to avoid another
rsync for that; especially if yet another timestep might be written during
that time.
Of course, as soon as a simulation finished all of this isn't an issue
anymore anyway.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/30#comment:4>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#30: Implement a "get" command
--------------------------+-------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Resolution: | Keywords:
--------------------------+-------------------------------------------------
Comment (by eschnett):
Stopping a simulation seems dangerous. I'd prefer something safer. Here is
one idea:
- Cactus writes a timestamp or some other unique identifier for each HDF5
file into a separate file.
- Simfactory checks this file before and after an rsync. If the time
stamps differ, the file has to be copied again.
Another, similar method would be to have Cactus generate checksums for the
HDF5 files. This is expensive, as the whole file would have to be
checksummed, unless there is an HDF5 facility for this. After copying the
file we compare checksums.
Both methods need to be combined with a method to avoid accessing a file
while it is being updated; renaming it to *.tmp is the standard Unix way
to do so.
Yet another way would be to ask Cactus to make copies of all output files.
Such copies are never changed after being created (only deleted if they
are outdated), and such can be safely copied.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/30#comment:3>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#30: Implement a "get" command
--------------------------+-------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Resolution: | Keywords:
--------------------------+-------------------------------------------------
Comment (by rhaas):
I believe option 2 will break if Cactus starts to write to a file while
simfactory is synching it. The OS will allow the rename and rsync will
happily continue reading from the now changing file, possibly mixing old
and new content. If Cactus finishes adding to the file before rsync
finishes and moves it from tmp to its original name, then simfactory would
not notice this (but rsync most likely would when it does its final check
on the file content).
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/30#comment:2>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#30: Implement a "get" command
--------------------------+-------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Resolution: | Keywords:
--------------------------+-------------------------------------------------
Comment (by hinder):
We will need some way to ensure that the retrieved data is in a consistent
state. Truncated ASCII files can be dealt with, though this is not ideal,
but partially-written HDF5 files cannot. This can be a serious problem if
several HDF5 files are being synced, as after each sync, there can be a
high probability that at least one of them is incompletely written. Some
options:
1. SimFactory (on the remote machine) writes a control file (either into
the simulation or somewhere else) which indicates to the simulation that
it should not open any new files for writing, and once it has closed any
currently open file as part of normal operation, it should record this
information in the control file, continue running, and only write new
files once the control file tells it to. This has the disadvantage of
locking the entire simulation for the duration of the transfer of the
active restart. For slow data transfers, this could be a significant
amount of time. This approach has the disadvantage that the different
files will not be in a consistent state; e.g. one output file may have the
current iteration but another may not.
2. Before writing a file, Cactus would move it to a new location
(file.tmp), and only move it back when it was fully written. SimFactory
would not sync tmp files. Any files which had been renamed to *.tmp in
the first pass would then be synced in a separate pass using their
original names, if they exist. Repeat until all files are synced. This
solution also does not maintain a consistent state across multiple files.
This does not require write access to the simulation directory, so could
also be used by collaborators who do not own the simulation.
3. Similar to (1), but only performed at the end of an iteration.
SimFactory would indicate to Cactus to pause the simulation at the end of
the current iteration, when all files are presumably valid on disk.
Cactus would indicate that the simulation had paused in a control file,
and SimFactory would then transfer the data, and unpause the simulation
when it was finished. This would guarantee that the synced data was in a
consistent state. We might want to have some mechanism to ensure that
simulations do not remain paused forever, perhaps by requiring simfactory
to update the control file periodically if it is still syncing.
All of the above apply only to the active restart. I think (3) is the
simplest and most robust. It is also the most expensive in SUs. The
control file location could be customisable, and placed somewhere that all
collaborators have write access.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/30#comment:1>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1257: known_architectures/linux doesn't deal with non-standard gcc version output
--------------------+-------------------------------------------------------
Reporter: knarf | Owner:
Type: defect | Status: new
Priority: minor | Milestone: ET_2013_05
Component: Cactus | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
Currently, known_architectures/linux only looks at the first line of '$F77
--version' to determine the version of a gnu-type fortran compiler.
However, on bluewaters this returns
{{{
/opt/cray/xt-asyncpe/5.16/bin/ftn: INFO: Compiling with
CRAYPE_COMPILE_TARGET=native.
GNU Fortran (GCC) 4.7.2 20120920 (Cray Inc.)
Copyright (C) 2012 Free Software Foundation, Inc.
}}}
The information the script is actually interested in is on the second
line, not the first. Thus, I propose the attached patch (grepping for
"GNU" before choosing the first line). I didn't find a problem with this
patch on other machines, but there might be in case $F77 --version does
not contain GNU in the line with the version.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1257>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1265: SimFactory does not support job chaining on supermuc
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: enhancement | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
-------------------------+--------------------------------------------------
Error: Machine supermuc currently does not support job chaining. Please
modify the submit command or submission script to support job chaining.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1265>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1264: only install those hdf5 utils that we find/could build
-----------------------------------+----------------------------------------
Reporter: rhaas | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: HDF5 |
-----------------------------------+----------------------------------------
the attached patch only tries to install those hdf5 utilities that were
eihter build by us or can be found in HDF5_DIR/bin since eg. a sytem
package provides them.
This helps with users that use the system provided hdf5 package but do not
install the hdf5-tools package (package names are debian names).
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1264>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1116: changing library options in optionlist does'nt trigger rebuild of depending
thorns
--------------------+-------------------------------------------------------
Reporter: knarf | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: Cactus | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
I first built Cactus using OpenMPI shipped with the thorn MPI. Then I
changed the optionlist to point to a local mpich2 installation and tried
to rebuild Cactus using the new optionlist (sim build
--thornlist=./thornlists/einsteintoolkit.th --optionlist=numrel-gcc.cfg).
The new options are used in config-info and on the link line, however,
thorns depending on MPI are not rebuilt, leading to link-failures because
these thorns have been built against OpenMPI.
Any change in the configuration of external libraries should lead to a
rebuilt of depending thorns.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1116>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#950: "Correct" Weyl Scalar symmetries
---------------------------------+------------------------------------------
Reporter: yosef@… | Type: defect
Status: new | Priority: major
Milestone: | Component: Other
Version: | Keywords:
---------------------------------+------------------------------------------
RotatingSymmetry180 and the ReflectionSymmetry thorns
define a tensor alias "weylscalars_real", but the symmetries given there
do not correspond to those of the usual re[psi0], im[psi0], re[psi1],
im[psi1],
etc scalars.
The Weyl symmetries as provided in RotatingSymmetry180
static int const weylparities[10][3] =
{{+1,+1,+1},
{-1,-1,-1},
{+1,+1,+1},
{-1,-1,-1},
{+1,+1,+1},
{-1,-1,-1},
{+1,+1,+1},
{-1,-1,-1},
{+1,+1,+1},
{-1,-1,-1}};
The symmetries for rpsi0, ipsi0, rpsi1, ipsi1, ..., rpsi4, ipsi4
static int const weylparities[10][3] =
{{+1,+1,+1}, /* rpsi0 */
{-1,-1,-1}, /* ipsi0 */
{+1,+1,-1}, /* rpsi1 */
{-1,-1,+1}, /* ipsi1 */
{+1,+1,+1}, /* rpsi2 */
{-1,-1,-1}, /* ipsi2 */
{+1,+1,-1}, /* rpsi3 */
{-1,-1,+1}, /* ipsi3 */
{+1,+1,+1}, /* rpsi4 */
{-1,-1,-1}}; /* ipsi4 */
attached is a diff between the current ET version of the Reflection and
RotatingSymmetry180 thorns and the "corrected" version.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/950>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit