#1193: CarpetReduce uses lsh for index calculations
--------------------+-------------------------------------------------------
Reporter: knarf | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone: ET_2013_05
Component: Carpet | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
CarpetReduce (e.g. reduce.cc:644) uses lsh to calculate GF indices. This
is wrong in case lsh!=ash. Either use ash or the Cactus macros (which
would probably be even better).
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1193>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1181: Parameter file parsing error message is hidden
----------------------+-----------------------------------------------------
Reporter: eschnett | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Cactus | Version:
Keywords: |
----------------------+-----------------------------------------------------
Parse errors in parameter files are hidden among other output, and are
thus difficult to spot. This example shows this:
{{{
cactus::cctk_itlast = 3 ;
ActiveThorns = "HDF5"
}}}
The problem is that the error message is output among the activation
messages. If the parameter file is long and the syntax error is in the
beginning, then several hundred lines may be output after the error
message, which makes it very difficult to spot it.
The parsing errors should be output in the same way as other parameter
errors, which are prefixed by WARNING etc.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1181>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1100: Correct backtrace generation in Carpet
----------------------+-----------------------------------------------------
Reporter: eschnett | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone:
Component: Carpet | Version:
Keywords: |
----------------------+-----------------------------------------------------
The file backtrace.cc in CarpetLib does not #include <cctk.h>; hence all
HAVE_BACKTRACE* macros are undefined, and only basic backtraces are
generated.
Correcting this is non-trivial, since the backtrace code is arcane, is
written in C, probably expects glibc, contains (I'm fairly certain) memory
allocation errors, and doesn't build e.g. on Mac OSX. The code also spends
an inordinate amount of time allocating and freeing string buffers, which
should be replaced by simply using C++ streams.
The backtrace code also probably requires a few more autoconf tests, so
that it can be disabled where it does not work.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1100>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1494: Allow skipping of MoL_PostStep and MoL_PseudoEvolutionBoundaries in
POSTRESTRICT
-------------------------+--------------------------------------------------
Reporter: hinder | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Cactus | Version: development version
Keywords: |
-------------------------+--------------------------------------------------
It is not always necessary to call MoL_PostStep and
MoL_PseudoEvolutionBoundaries in POSTRESTRICT, and it can introduce a
performance penalty. The main reason for these calls is that restriction
does not fill (outer or symmetry) boundary points, and this is usually
done in MoL_PostStep. MoL_PseudoEvolutionBoundaries also sets boundary
conditions. However, if restriction does not modify boundary points, for
example in the case that boundary points are always far from refined
regions, there is no reason to apply boundary conditions (e.g. by calling
MoL_PostStep) after restriction.
Eventually, Carpet and MoL should be modified to determine automatically
whether the BCs need to be applied, but until that is implemented, the
attached patch provides parameters for careful users to optimise their
simulations in the case where this is safe to do.
Additionally, recalculations performed in MoL_PostStep may replace more
accurate restricted values computed on finer grids, leading to a loss of
accuracy.
OK to commit?
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1494>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#914: Don't use fork()
-----------------------------------+----------------------------------------
Reporter: eschnett | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: |
-----------------------------------+----------------------------------------
It seems that it is in many cases not safe to call fork() in MPI
applications. This page <http://www.open-mpi.de/faq/?category=openfabrics
#ofa-fork> has some information. The upshot seems to be:
- In many (most) cases, one can call system() or popen() to execute
external processes while waiting for them.
- It is generally not safe to call fork() to execute a certain task in the
background. However, it should be possible to use threads in this case.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/914>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#260: produce map of ET users
-------------------------------------+--------------------------------------
Reporter: knarf | Owner:
Type: enhancement | Status: new
Priority: optional | Milestone:
Component: EinsteinToolkit website | Version:
Keywords: |
-------------------------------------+--------------------------------------
It would be nice to produce an (autmatically generated) map of the
locations of ET users
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/260>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1542: Define variants of CCTK_FullName and CCTK_GroupName that don't require
calling free
-------------------------+--------------------------------------------------
Reporter: eschnett | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Cactus | Version: development version
Keywords: |
-------------------------+--------------------------------------------------
'nuff said.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1542>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1581: Clean up ET web site menu
-------------------------------------+--------------------------------------
Reporter: eschnett | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: EinsteinToolkit website | Version: development version
Keywords: |
-------------------------------------+--------------------------------------
The menu entries "support" and "issue tracker" should be separated from
"wiki", "blog", and "seminars". The first two are about the code, the
latter two about the community. I suggest to move the first two to the
"download" section that currently has no sub-menus.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1581>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1547: issues with stampede
----------------------------+-----------------------------------------------
Reporter: jchsma@… | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Other | Version: ET_2013_11
Keywords: |
----------------------------+-----------------------------------------------
Over the past few months, I have been using RIT's LazEv code with only
minor hiccups on stampede (particularly, the unreproducible 'dapl_conn_rc'
crashes that I'm sure other stampede users are familiar with). This
checkout was of the previous release, ET_2013_05, and compiled with Intel
MPI. Most of the jobs I ran took advantage of some symmetry, and I was
able to run on 12-16 nodes at about 50-60% memory usage.
After the sync issue was backported, I checked out the new release,
ET_2013_11 and immediately ran into problems. The first issue was with
the run performance and LoopControl, which, with the mailing lists help,
we sorted out. The second was with crashes and checkpointing. With both
Intel MPI and MVAPICH2 configurations, the code would hang (~50% of the
time) when dumping checkpoints, and 100% when dumping a termination
checkpoint. Further, the crashes seem more frequent, and I couldn't get a
simulation to run for a full 24 hours without crashing (either by stalling
on checkpointing or otherwise).
So, I checked out a clean version of the toolkit, with only toolkit
thorns, and removed any thorns specific to RIT. I compiled with both the
Intel MPI and MVAPICH2 configurations in simfactory.
In both cases, I can run the 'qc0-mclachlan.par' file to completion with
no issues. So I edited the qc0 parfile to update the grid, remove the
symmetries, and update the initial data to match my test parameter file.
I ran the job on 20 nodes, and with either configuration, I was not
successful in running the job to completion on any of my numerous
attempts. Intel MPI runs die with the standard unhelpful "dapl_conn_rc"
error at random times in the evolution, and the MVAPICH2 dies with:
[c431-903.stampede.tacc.utexas.edu:mpispawn_7][readline] Unexpected End-
Of-File on file descriptor 6. MPI process died?
[c431-903.stampede.tacc.utexas.edu:mpispawn_7][mtpmi_processops] Error
while reading PMI socket. MPI process died?
[c431-903.stampede.tacc.utexas.edu:mpispawn_7][child_handler] MPI process
(rank: 15, pid: 106620) terminated with signal 9 -> abort job
[c429-501.stampede.tacc.utexas.edu:mpirun_rsh][process_mpispawn_connection]
mpispawn_7 from node c431-903 aborted: Error while reading a PMI socket
(4)
The IMPI jobs died with the same dapl_conn_rc error at run times of 2
hours, 8 hours, and 21 hours. I also had one job that hung and did not
exit until it was killed by the queue manager. The MVAPICH2 jobs died at
around 3 hours and 8 hours with the error above.
We've been in contact with TACC and they said it was a Cactus issue, so I
am sending this report.
Attached is the parameter file I used for the tests. They should work
with a stock ET_2013_11 checkout.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1547>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#980: Remove support for the flesh-based MPI mechanism
-------------------------+--------------------------------------------------
Reporter: hinder | Owner:
Type: enhancement | Status: new
Priority: major | Milestone:
Component: Cactus | Version:
Keywords: |
-------------------------+--------------------------------------------------
The flesh-based MPI mechanism has just been deprecated in favour of
ExternalLibraries/MPI. Even though the latter is new, it is probably a
good idea to completely disable the old mechanism since having any mixture
of thorns/optionlists using the old and new mechanisms is completely
untested and will likely lead to problems and confusion. It's better to
give a useful fatal error message if someone still specifies "MPI = " in
their optionlist than to have things break in other weird and wonderful
ways.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/980>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit