#1327: Delay subsequent restarts in the case of certain problems
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: enhancement | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
-------------------------+--------------------------------------------------
If one restart of a simulation exits abnormally, e.g. due to some
transient problem on a cluster, all subsequent restarts might also run
into the same problem. If we can distinguish between terminations due to
internal (i.e. numerical or code-related) problems and external (MPI
errors, filesystem issues) problems, we can do different things for each.
Possible actions could be:
1. Continue as normal with the next restart;
2. Delay the next restart for a few hours, in the hope that the transient
cluster problems are resolved;
3. Hold the next restart and notify the user by email that an
unrecoverable error has occurred.
These could be communicated by exit codes (whether through official
methods, or through an exit code file). Distinguishing between 2 and 3
could be achieved by regular expression matching on the standard output or
standard error file. This would make the mechanism independent of Cactus.
So Cactus would only have to say "good" or "bad", and SimFactory could
then decide if "bad" meant to delay or hold based on some logic in its
machine database.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1327>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#322: SimFactory metadata deleted by periodic filesystem purges
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Keywords: |
-------------------------+--------------------------------------------------
Production filesystems are subject to periodic purges (typically on the
order of weeks or months) where data which has not been accessed recently
is deleted. This means that it is possible for some restarts of very
long-running simulations to be deleted by the system. This can be
addressed by an automated archiving system, but such a system does not
address the problem that the simulation metadata directory (currently
called SIMFACTORY) and any restarts which have not been run yet, will also
be purged. This would make it impossible to submit future restarts and
limits the number of chained restarts you can submit to the purge time of
the system.
One possibility to solve this problem would be to store a backup, or
"shadow" copy of all the simulation metadata in a non-volatile location.
This could be the user's home directory, or a "work" directory which is
not purged. The details would need to be worked out.
This is not a serious issue yet.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/322>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1323: remove make.code.deps from GRHydro
-----------------------------------+----------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: GRHydro |
-----------------------------------+----------------------------------------
GRHydro contains a file make.code.deps which tries to list the
dependencies among its files due to Fortran modules. Naturally the file is
out of date. Reading the documentation this should be harmless as
make.code.deps should only give additional dependency information. However
I just had a case where removing the file (and -cleandeps) helped. The
actual fix could thus have been the cleandeps as well.
make.code.deps is not required in this case since Cactus can track module
dependency on its own (see
http://einsteintoolkit.org/documentation/UsersGuide/UsersGuidech9.html#x13-…
and
http://einsteintoolkit.org/documentation/UsersGuide/UsersGuidech9.html#x13-…).
This patch removes the file.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1323>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1325: McLachlan should not checkpoint the constraint variables if they have only
one timelevel
-----------------------------------+----------------------------------------
Reporter: hinder | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: McLachlan |
-----------------------------------+----------------------------------------
ML_BSSN_Helper currently checkpoints the constraint variables even if they
only have one timelevel. This is because it incorrectly assumes that the
number of timelevels for these variables is given by the "timelevels"
parameter, when in fact it is given by the "other_timelevels" parameter.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1325>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1310: HydroBase should initialize excision mask before MoL_PostStep since it is
used in MoL_PostStep
----------------------------------------+-----------------------------------
Reporter: reisswig@… | Owner: reisswig@…
Type: defect | Status: new
Priority: critical | Milestone: ET_2013_05
Component: EinsteinToolkit thorn | Version: development version
Keywords: HydroBase |
----------------------------------------+-----------------------------------
Some routines in GRHydro query whether a point is excised using
hydrobase's excision mask. HydroBase initializes the excision mask to zero
in certain groups.
However, this initialization should happen before any other code reads the
excision mask.
Since GRHydro reads the excision mask in Con2Prim, which is scheduled in
MoL_PostStep, the mask should be set before MoL_PostStep!
Most notably, this is required in post_recover_variables (since the mask
is not checkpointed and thus we critically rely on a correct mask
initialzation), and also in post_regrid, where new points must be
correctly initialized.
I attach a simple patch that fixes this issue.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1310>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1322: CarpetRegrid2 movement_threshold
-----------------------------------+----------------------------------------
Reporter: vassilios.mewes@… | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version: ET_2012_11
Keywords: CarpetRegrid2 |
-----------------------------------+----------------------------------------
i have got a question about the CarpetRegrid2 process of regridding when
supplying the thorn with information from puncture tracker...
in a recent simulation of BH+torus system i did, the BH was moving
significantly further than the movement_threshold i supplied in the par
file, but Carpet never regridded...it seems to me that the new center is
set at a frequency of regrid_every, but shouldn't it be reset only when
the center has moved further than the movement_threshold parameter? as it
seems to me at the moment, i would have to know the movement of the BH a
priori in order to chose a sensible regridding frequency, because if the
BH moves too little between regrid_every, it will eventually move out of
the finest refinement level..
furthermore, the regridding frequency might change during the run because
the BH is actually accelerating...
i am doing something wrong here or setting false parameters?
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1322>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#591: Add harmonic shift to McLachlan
-----------------------------------+----------------------------------------
Reporter: hinder | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: McLachlan |
-----------------------------------+----------------------------------------
The attached patch adds a harmonic shift condition to McLachlan. It does
this by introducing a new real-valued parameter harmonicShift. This
should be set to 0 for gamma-driver shift (the default, and the existing
behaviour), and to 1 for harmonic shift. This is useful for code-
correctness tests with the shifted gauge wave which is an exact solution
of the Einstein equations in harmonic gauge. The harmonic shift equation
has been tested with the shifted gauge wave exact solution and yields
convergence to the exact solution.
The current way that gauge conditions is handled in McLachlan is not very
elegant, and I don't think this patch should be applied as-is. I am
putting it here for anyone who might find it useful.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/591>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1289: Harmonic shift in McLachlan only works for conformalMethod = 1 (W method)
-----------------------------------+----------------------------------------
Reporter: hinder | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: McLachlan |
-----------------------------------+----------------------------------------
The harmonic shift implemented in #591 is missing the case discrimination
for the conformalMethod parameter; the way it is currently coded is
correct for conformalMethod = 1 (W method) but not for conformalMethod = 0
(phi method). My original version of the code had the two variants but I
removed this at some point while testing and my submitted patch contained
only conformalMethod = 1. In my original notes, I have this
{{{
alpha^2 em4phi (gtu[ua,
uk] (PD[alpha, lk]/alpha +
2 IfThen[conformalMethod, -1/(2 phi), 1] PD[phi, lk]) -
gtu[ua, ul] gtu[uj, um] PD[gt[ll, lm], lj])
}}}
which should be checked again before committing.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1289>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1321: Collisions in executable cache directory
------------------------+---------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
If I have two Cactus trees both using the same configuration names and the
same simulation directory, I believe that the executable cache in
simulations/CACHE will suffer from collisions between the cached
executables in the different Cactus trees. A solution would be to use a
unique name for the executable, just as is done when purging simulations
to the TRASH directory.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1321>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1304: slow build
----------------------------------+-----------------------------------------
Reporter: srichers@… | Owner: sbrandt
Type: defect | Status: new
Priority: major | Milestone:
Component: Mojave | Version:
Keywords: |
----------------------------------+-----------------------------------------
There is a significant lag when using the Mojave 'Build' command. It is
not noticeable for the WaveToy demo, but for the Einstein Toolkit, eclipse
freezes for about 30 seconds before the console shows build progress.
Following a completed build ("done." printed in the console), the progress
bar in eclipse remains static at one percentage for another ~5 minutes,
preventing another build command from executing.
Fresh (yesterday) install of eclipse and Mojave. ET downloaded through
Mojave wizard using development thornlist on the ET website.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1304>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit