#1335: SimFactory should abort if there are no checkpoint files when submitting an
existing configuration
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: enhancement | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
-------------------------+--------------------------------------------------
When submitting a simulation which already contains at least one restart,
simfactory should abort if there are no checkpoint files available. This
likely means that something went wrong. Starting the simulation again is
always the wrong thing to do in this case, as it will waste CPU time and
might go unnoticed.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1335>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1329: GetComponents is not asking for svn password
---------------------------+------------------------------------------------
Reporter: knarf | Owner: eric9
Type: defect | Status: new
Priority: minor | Milestone: ET_2013_05
Component: GetComponents | Version: development version
Keywords: |
---------------------------+------------------------------------------------
GetComponents uses the --non-interactive option to svn even for an initial
checkout, which, even if needed, doesn't ask for an svn password. The
effect is that if svn passwords are not already stored, a checkout using
-noa fails on every repository that has to use passwords.
That means that we probably have to remove the --non-interactive option
from svn checkout commands. GetComponents has to let svn handle this.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1329>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1332: Prolongation operators are only parallelized in one direction
-------------------------+--------------------------------------------------
Reporter: knarf | Owner: eschnett
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Carpet | Version: development version
Keywords: |
-------------------------+--------------------------------------------------
Currently, Carpet's prolongation operators are only parallelized in one
direction (the largest). This happens before calling the actual operator
in call_operator() (CarpetLib). With increasing core(thread)-counts this
is a problem. There are hardly enough points in any single direction to
make this efficient.
Wouldn't it be much better (in terms of openmp efficiency) to let the
operators handle openmp-parallelization themselves? I currently see quite
a large overhead for real-world parfiles with 32 threads just because of
this.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1332>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1331: www.einsteintoolkit.org leads to apache test page
-------------------------------------+--------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit website | Version:
Keywords: |
-------------------------------------+--------------------------------------
entering http://www.einsteintoolkit.org instead of
http://einsteintoolkit.org brings me to the apache test page. Since this
might be a common mistake, we might want to put a redirection page at the
incorrect page.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1331>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1330: hwloc might build against cuda, in which case Cactus needs to link against
it
-----------------------------------+----------------------------------------
Reporter: knarf | Owner:
Type: defect | Status: new
Priority: major | Milestone: ET_2013_05
Component: EinsteinToolkit thorn | Version: development version
Keywords: |
-----------------------------------+----------------------------------------
Currently, the hwloc pkg-config does not work in some cases, and the
'reasonable guess' in configure.sh doesn't as well. The main problem seems
to be that pkg-config is called before PKG_CONFIG_PATH is set.
The attached patch fixes this and allows me to build on a system where
cuda libraries are installed (and the hwloc autoconf finds them and links
against them), but Cactus isn't configured for cuda. Without the patch
linking Cactus fails because of missing cuda libraries.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1330>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1327: Delay subsequent restarts in the case of certain problems
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: enhancement | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
-------------------------+--------------------------------------------------
If one restart of a simulation exits abnormally, e.g. due to some
transient problem on a cluster, all subsequent restarts might also run
into the same problem. If we can distinguish between terminations due to
internal (i.e. numerical or code-related) problems and external (MPI
errors, filesystem issues) problems, we can do different things for each.
Possible actions could be:
1. Continue as normal with the next restart;
2. Delay the next restart for a few hours, in the hope that the transient
cluster problems are resolved;
3. Hold the next restart and notify the user by email that an
unrecoverable error has occurred.
These could be communicated by exit codes (whether through official
methods, or through an exit code file). Distinguishing between 2 and 3
could be achieved by regular expression matching on the standard output or
standard error file. This would make the mechanism independent of Cactus.
So Cactus would only have to say "good" or "bad", and SimFactory could
then decide if "bad" meant to delay or hold based on some logic in its
machine database.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1327>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#322: SimFactory metadata deleted by periodic filesystem purges
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Keywords: |
-------------------------+--------------------------------------------------
Production filesystems are subject to periodic purges (typically on the
order of weeks or months) where data which has not been accessed recently
is deleted. This means that it is possible for some restarts of very
long-running simulations to be deleted by the system. This can be
addressed by an automated archiving system, but such a system does not
address the problem that the simulation metadata directory (currently
called SIMFACTORY) and any restarts which have not been run yet, will also
be purged. This would make it impossible to submit future restarts and
limits the number of chained restarts you can submit to the purge time of
the system.
One possibility to solve this problem would be to store a backup, or
"shadow" copy of all the simulation metadata in a non-volatile location.
This could be the user's home directory, or a "work" directory which is
not purged. The details would need to be worked out.
This is not a serious issue yet.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/322>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1323: remove make.code.deps from GRHydro
-----------------------------------+----------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: GRHydro |
-----------------------------------+----------------------------------------
GRHydro contains a file make.code.deps which tries to list the
dependencies among its files due to Fortran modules. Naturally the file is
out of date. Reading the documentation this should be harmless as
make.code.deps should only give additional dependency information. However
I just had a case where removing the file (and -cleandeps) helped. The
actual fix could thus have been the cleandeps as well.
make.code.deps is not required in this case since Cactus can track module
dependency on its own (see
http://einsteintoolkit.org/documentation/UsersGuide/UsersGuidech9.html#x13-…
and
http://einsteintoolkit.org/documentation/UsersGuide/UsersGuidech9.html#x13-…).
This patch removes the file.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1323>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1325: McLachlan should not checkpoint the constraint variables if they have only
one timelevel
-----------------------------------+----------------------------------------
Reporter: hinder | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: McLachlan |
-----------------------------------+----------------------------------------
ML_BSSN_Helper currently checkpoints the constraint variables even if they
only have one timelevel. This is because it incorrectly assumes that the
number of timelevels for these variables is given by the "timelevels"
parameter, when in fact it is given by the "other_timelevels" parameter.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1325>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1310: HydroBase should initialize excision mask before MoL_PostStep since it is
used in MoL_PostStep
----------------------------------------+-----------------------------------
Reporter: reisswig@… | Owner: reisswig@…
Type: defect | Status: new
Priority: critical | Milestone: ET_2013_05
Component: EinsteinToolkit thorn | Version: development version
Keywords: HydroBase |
----------------------------------------+-----------------------------------
Some routines in GRHydro query whether a point is excised using
hydrobase's excision mask. HydroBase initializes the excision mask to zero
in certain groups.
However, this initialization should happen before any other code reads the
excision mask.
Since GRHydro reads the excision mask in Con2Prim, which is scheduled in
MoL_PostStep, the mask should be set before MoL_PostStep!
Most notably, this is required in post_recover_variables (since the mask
is not checkpointed and thus we critically rely on a correct mask
initialzation), and also in post_regrid, where new points must be
correctly initialized.
I attach a simple patch that fixes this issue.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1310>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit