#1283: Missing data in HDF5 files
--------------------+-------------------------------------------------------
Reporter: hinder | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Cactus | Version:
Keywords: |
--------------------+-------------------------------------------------------
If a simulation runs out of disk space while writing an HDF5 file, the
simulation will terminate. The hdf5 file being written might then be
corrupt, and all data from it may be irretrievable. In that case,
restarting from the last-written checkpoint file may leave a "gap" in the
data corresponding to the period between the start of the failed restart
and the last checkpoint file written.
Steps to reproduce:
• Start a simulation which checkpoints periodically and consists
of several restarts
• Keep all checkpoint files
• Restart 0000 completes successfully and checkpoints at iteration
i1
• Restart 0001 checkpoints once after some evolution at iteration
i2
• Restart 0001 terminates abnormally while writing an HDF5 output
file at iteration i3
• The output file is corrupted and nonrecoverable, so there is no
data from iteration i1 to iteration i3
• Restart 0002 starts at iteration i2 as this is the last
checkpoint available
• The simulation continues until the end, but the data from the
corrupted HDF5 file between iteration i1 and i2 is lost
Possible solutions:
1. Write HDF5 files safely, e.g. by first copying the file to a
new temporary file, performing the write, then atomically moving the
temporary file over the original file. The original file would then
remain in the event of a crash while writing the new file. This could be
very expensive for 3D output files.
2. Start a new set of HDF5 files after each checkpoint. This
seems to be the most efficient and simplest solution, but requires readers
of HDF5 files to be modified to take it into account.
3. Check the consistency of all HDF5 files in the previous
restart(s) on recovery, and recover from the latest checkpoint file for
which all previous HDF5 files are valid. We could use code to check the
HDF5 file, or some other flagging mechanism to indicate that HDF5 writes
were completed successfully; e.g. we could rename the HDF5 file to .tmp
during writes, and rename it back after a successful write. This is
complex and requires Cactus or simfactory to look into previous restarts.
It also only applies to HDF5 files, and requires breaking several
abstraction barriers.
4. Wait for HDF5 journalling support. As far as I know only
metadata journalling is planned, which is probably not enough, and in any
case, they are not actively working on the next version of HDF5 at the
moment due to lack of funding.
5. Checkpoint only on termination of the simulation
In reality, we do not keep all checkpoint files. I usually keep just the
last checkpoint file. I believe that a Cactus simulation will only delete
checkpoint files which it has itself written, which means that there will
generally be one checkpoint file kept per restart; the last one written.
This means that you can always recover from the above situation by
rerunning the restart during which the problem occurred. However, keeping
one checkpoint file per restart is a problem in itself, and we should fix
this as well, which would then mean the potential for losing data in the
case of an interrupted write operation.
Thoughts?
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1283>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1193: CarpetReduce uses lsh for index calculations
--------------------+-------------------------------------------------------
Reporter: knarf | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone: ET_2013_05
Component: Carpet | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
CarpetReduce (e.g. reduce.cc:644) uses lsh to calculate GF indices. This
is wrong in case lsh!=ash. Either use ash or the Cactus macros (which
would probably be even better).
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1193>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1181: Parameter file parsing error message is hidden
----------------------+-----------------------------------------------------
Reporter: eschnett | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Cactus | Version:
Keywords: |
----------------------+-----------------------------------------------------
Parse errors in parameter files are hidden among other output, and are
thus difficult to spot. This example shows this:
{{{
cactus::cctk_itlast = 3 ;
ActiveThorns = "HDF5"
}}}
The problem is that the error message is output among the activation
messages. If the parameter file is long and the syntax error is in the
beginning, then several hundred lines may be output after the error
message, which makes it very difficult to spot it.
The parsing errors should be output in the same way as other parameter
errors, which are prefixed by WARNING etc.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1181>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1100: Correct backtrace generation in Carpet
----------------------+-----------------------------------------------------
Reporter: eschnett | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone:
Component: Carpet | Version:
Keywords: |
----------------------+-----------------------------------------------------
The file backtrace.cc in CarpetLib does not #include <cctk.h>; hence all
HAVE_BACKTRACE* macros are undefined, and only basic backtraces are
generated.
Correcting this is non-trivial, since the backtrace code is arcane, is
written in C, probably expects glibc, contains (I'm fairly certain) memory
allocation errors, and doesn't build e.g. on Mac OSX. The code also spends
an inordinate amount of time allocating and freeing string buffers, which
should be replaced by simply using C++ streams.
The backtrace code also probably requires a few more autoconf tests, so
that it can be disabled where it does not work.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1100>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1494: Allow skipping of MoL_PostStep and MoL_PseudoEvolutionBoundaries in
POSTRESTRICT
-------------------------+--------------------------------------------------
Reporter: hinder | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Cactus | Version: development version
Keywords: |
-------------------------+--------------------------------------------------
It is not always necessary to call MoL_PostStep and
MoL_PseudoEvolutionBoundaries in POSTRESTRICT, and it can introduce a
performance penalty. The main reason for these calls is that restriction
does not fill (outer or symmetry) boundary points, and this is usually
done in MoL_PostStep. MoL_PseudoEvolutionBoundaries also sets boundary
conditions. However, if restriction does not modify boundary points, for
example in the case that boundary points are always far from refined
regions, there is no reason to apply boundary conditions (e.g. by calling
MoL_PostStep) after restriction.
Eventually, Carpet and MoL should be modified to determine automatically
whether the BCs need to be applied, but until that is implemented, the
attached patch provides parameters for careful users to optimise their
simulations in the case where this is safe to do.
Additionally, recalculations performed in MoL_PostStep may replace more
accurate restricted values computed on finer grids, leading to a loss of
accuracy.
OK to commit?
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1494>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#914: Don't use fork()
-----------------------------------+----------------------------------------
Reporter: eschnett | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: |
-----------------------------------+----------------------------------------
It seems that it is in many cases not safe to call fork() in MPI
applications. This page <http://www.open-mpi.de/faq/?category=openfabrics
#ofa-fork> has some information. The upshot seems to be:
- In many (most) cases, one can call system() or popen() to execute
external processes while waiting for them.
- It is generally not safe to call fork() to execute a certain task in the
background. However, it should be possible to use threads in this case.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/914>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#260: produce map of ET users
-------------------------------------+--------------------------------------
Reporter: knarf | Owner:
Type: enhancement | Status: new
Priority: optional | Milestone:
Component: EinsteinToolkit website | Version:
Keywords: |
-------------------------------------+--------------------------------------
It would be nice to produce an (autmatically generated) map of the
locations of ET users
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/260>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1781: Outflow: wrong column labels in 2D output
-----------------------------------+----------------------------------------
Reporter: dradice@… | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version: development version
Keywords: |
-----------------------------------+----------------------------------------
This is a trivial fix (I am attaching a patch)
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1781>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1542: Define variants of CCTK_FullName and CCTK_GroupName that don't require
calling free
-------------------------+--------------------------------------------------
Reporter: eschnett | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Cactus | Version: development version
Keywords: |
-------------------------+--------------------------------------------------
'nuff said.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1542>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1581: Clean up ET web site menu
-------------------------------------+--------------------------------------
Reporter: eschnett | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: EinsteinToolkit website | Version: development version
Keywords: |
-------------------------------------+--------------------------------------
The menu entries "support" and "issue tracker" should be separated from
"wiki", "blog", and "seminars". The first two are about the code, the
latter two about the community. I suggest to move the first two to the
"download" section that currently has no sub-menus.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1581>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit