#1262: simfactory doesn't know how to handle 'old' simulations
------------------------+---------------------------------------------------
Reporter: knarf | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone: ET_2013_05
Component: SimFactory | Version: development version
Keywords: |
------------------------+---------------------------------------------------
Trying to restart a simulation after an update of simfactory I get
{{{
$ sim submit A3_oct --procs=256
Assigned restart id: 2
Traceback (most recent call last):
File "/home/knarf/utils/simfactory/bin/../lib/sim.py", line 148, in
<module>
main()
File "/home/knarf/utils/simfactory/bin/../lib/sim.py", line 144, in main
CommandDispatch()
File "/home/knarf/utils/simfactory/bin/../lib/sim.py", line 106, in
CommandDispatch
module.main()
File "/home/knarf/utils/simfactory/lib/sim-manage.py", line 397, in main
CommandDispatch()
File "/home/knarf/utils/simfactory/lib/sim-manage.py", line 376, in
CommandDispatch
exec("command_%s()" % command)
File "<string>", line 1, in <module>
File "/home/knarf/utils/simfactory/lib/sim-manage.py", line 267, in
command_submit
restart.userSubmit(simulationName)
File "/home/knarf/utils/simfactory/lib/simrestart.py", line 353, in
userSubmit
self.submit(submitScript)
File "/home/knarf/utils/simfactory/lib/simrestart.py", line 590, in
submit
(nodes, ppn_used, procs, ppn, procs_requested, num_procs, num_threads,
num_smt) = simlib.GetProcs(existingProperties)
File "/home/knarf/utils/simfactory/lib/simlib.py", line 834, in GetProcs
num_smt = existingProperties.numsmt
AttributeError: SimProperties instance has no attribute 'numsmt'
}}}
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1262>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1273: hdf5 index files created for checkpoints not deleted
-----------------------------------------+----------------------------------
Reporter: wolfgang.kastaun@… | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: Carpet | Version: ET_2012_11
Keywords: |
-----------------------------------------+----------------------------------
I recently evolved with activated hdf5 index file creation, using the
Oersted release. I notice that also for checkpoint files the index files
are saved, but in contrast to the checkpoint files themselves, the old
index files are never deleted. I had more than 25000 checkpoint index
files accumulated in my simulation directory !
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1273>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1269: use XARGS configuration variable
--------------------+-------------------------------------------------------
Reporter: knarf | Owner: knarf
Type: defect | Status: new
Priority: minor | Milestone: ET_2013_05
Component: Cactus | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
The attached patch let's Cactus use a define variable XARGS, much like
others (e.g., TAR). This is necessary on systems where the default xargs
program cannot be used (e.g., pandora.hpc.lsu.edu). The patch also let's
the flesh use this variable, making it compile on these machines.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1269>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1139: Adding PeriodicCarpet to the Einstein Toolkit
-------------------------+--------------------------------------------------
Reporter: bentivegna | Owner: eschnett
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Carpet | Version: development version
Keywords: |
-------------------------+--------------------------------------------------
Old thorn Periodic does not behave correctly with AMR patches that overlap
the periodic boundaries (see ticket #694). A significantly different
mechanism to handle has been implemented, by Erik, in
LSUThorns/PeriodicCarpet. This thorn should be added to the Einstein
Toolkit, and slowly replace Periodic.
Before introducing this thorn to the ET, we need at least one testcase
with AMR.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1139>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1268: Cross-references in reference manual should be easier
-------------------------+--------------------------------------------------
Reporter: eschnett | Owner:
Type: enhancement | Status: new
Priority: major | Milestone:
Component: Cactus | Version:
Keywords: |
-------------------------+--------------------------------------------------
The cross references in the reference manual look like:
{{{
\begin{SeeAlso2}{CCTK\_VError}{CCTK-VError}
prints an error message with a variable argument list
\end{SeeAlso2}
}}}
This is cumbersome, because the function description needs to be repeated.
Instead, a simple pointer to the function should suffice. If this is not
possible directly in latex, then we may want to use a different mechanism
(e.g. Doxygen?).
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1268>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1186: Provide a version of CCTK_WARN which never returns
-------------------------+--------------------------------------------------
Reporter: hinder | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Cactus | Version:
Keywords: |
-------------------------+--------------------------------------------------
CCTK_WARN(0,...) always causes a fatal abort, but the compiler does not
know this. Hence, it incorrectly warns that variables might be used
without being set because it doesn't know that the only path that doesn't
set them terminates the program. I propose introducing a new function
CCTK_WARN_ABORT(...) which calls CCTK_WARN(0,...) and is declared with the
attribute noreturn (http://gcc.gnu.org/onlinedocs/gcc-3.2/gcc/Function-
Attributes.html) so that the compiler knows that it can never return, and
does not make incorrect conclusions when computing variable dependencies
(CCTK_Abort is already defined without an error message).
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1186>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1255: change default of IOUtil::out_dir from "." to "$parfile"
-------------------------+--------------------------------------------------
Reporter: knarf | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Cactus | Version: development version
Keywords: |
-------------------------+--------------------------------------------------
Currently the default output directory for Cactus output is ".", something
that is hardly every used by anybody. We should strive to have parameter
defaults that work in most circumstances and that are used by most people.
"." is save, but not at all the most wanted choice. "$parfile" is another
sensible choice, should be as save as "." and is used by more (>0) users
than their default. Thus, I propose to test to change the default. Before
getting into this (unknown) amount of work, let me use this ticket to ask
for objections first.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1255>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#30: Implement a "get" command
--------------------------+-------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Resolution: | Keywords:
--------------------------+-------------------------------------------------
Comment (by hinder):
I think that would already be a good solution. Yes, when I have found an
odd error, I find that h5ls fails, or gives no datasets. This may not be
100% reliable, but it's better than nothing. Yes, you will never get all
the files consistent unless you lock the simulation or have a filesystem
snapshot. I agree that temporal consistency is much less important than
avoiding corrupt hdf5 files; I already have code to handle temporal
inconsistency (two BH_diagnostics output files having different numbers of
iterations, etc).
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/30#comment:7>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#30: Implement a "get" command
--------------------------+-------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Resolution: | Keywords:
--------------------------+-------------------------------------------------
Comment (by knarf):
I don't have a corrupted hdf5 file handy, so I have to ask: Is there a
quick way to check if such a file is corrupt, maybe an 'h5ls'? If so, this
would make a selected resync quite easy.
Especially for high-frequency output you might come into trouble with
timestamps if your rsync takes longer than one of the output frequency
steps (and the simulation doesn't stop, which I really don't want to
have). You'll never get all files consistently because the moment you copy
the last the first has already been updated again.
All I care about would be to have all files non-corrupt. Retrying (only
the corrupt files) until this is reached sounds like a possible solution
to me, and would be an incentive to use simfactory over the otherwise also
simple, manual rsync.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/30#comment:6>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#30: Implement a "get" command
--------------------------+-------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Resolution: | Keywords:
--------------------------+-------------------------------------------------
Comment (by hinder):
I have been working like this for the last few years, and the occasional
corrupted HDF5 has not been a problem. However, recently I have started
to run simulations with many more files and more frequent output, and I
sometimes get into the situation where every time I sync the simulation,
at least one of the files is corrupt, as they are written very frequently,
and if they are copied before the write finishes, the copy can be corrupt.
A corrupt HDF5 file cannot be used; you cannot read the previous datasets,
you cannot even list the datasets. You just have to sync again and hope
for the best. I also agree that stopping the simulation is to be avoided
if possible. In the far future, we would probably be able to create a
filesystem-level snapshot of the simulation at the end of each iteration.
That would ensure a consistent state which could be easily copied. But
filesystems and OSes are not up to that yet.
Re: writing a time stamp to a separate file: The timestamp on the file
won't be updated until the file is closed.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/30#comment:5>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit