#1547: issues with stampede
----------------------------+-----------------------------------------------
Reporter: jchsma@… | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Other | Version: ET_2013_11
Keywords: |
----------------------------+-----------------------------------------------
Over the past few months, I have been using RIT's LazEv code with only
minor hiccups on stampede (particularly, the unreproducible 'dapl_conn_rc'
crashes that I'm sure other stampede users are familiar with). This
checkout was of the previous release, ET_2013_05, and compiled with Intel
MPI. Most of the jobs I ran took advantage of some symmetry, and I was
able to run on 12-16 nodes at about 50-60% memory usage.
After the sync issue was backported, I checked out the new release,
ET_2013_11 and immediately ran into problems. The first issue was with
the run performance and LoopControl, which, with the mailing lists help,
we sorted out. The second was with crashes and checkpointing. With both
Intel MPI and MVAPICH2 configurations, the code would hang (~50% of the
time) when dumping checkpoints, and 100% when dumping a termination
checkpoint. Further, the crashes seem more frequent, and I couldn't get a
simulation to run for a full 24 hours without crashing (either by stalling
on checkpointing or otherwise).
So, I checked out a clean version of the toolkit, with only toolkit
thorns, and removed any thorns specific to RIT. I compiled with both the
Intel MPI and MVAPICH2 configurations in simfactory.
In both cases, I can run the 'qc0-mclachlan.par' file to completion with
no issues. So I edited the qc0 parfile to update the grid, remove the
symmetries, and update the initial data to match my test parameter file.
I ran the job on 20 nodes, and with either configuration, I was not
successful in running the job to completion on any of my numerous
attempts. Intel MPI runs die with the standard unhelpful "dapl_conn_rc"
error at random times in the evolution, and the MVAPICH2 dies with:
[c431-903.stampede.tacc.utexas.edu:mpispawn_7][readline] Unexpected End-
Of-File on file descriptor 6. MPI process died?
[c431-903.stampede.tacc.utexas.edu:mpispawn_7][mtpmi_processops] Error
while reading PMI socket. MPI process died?
[c431-903.stampede.tacc.utexas.edu:mpispawn_7][child_handler] MPI process
(rank: 15, pid: 106620) terminated with signal 9 -> abort job
[c429-501.stampede.tacc.utexas.edu:mpirun_rsh][process_mpispawn_connection]
mpispawn_7 from node c431-903 aborted: Error while reading a PMI socket
(4)
The IMPI jobs died with the same dapl_conn_rc error at run times of 2
hours, 8 hours, and 21 hours. I also had one job that hung and did not
exit until it was killed by the queue manager. The MVAPICH2 jobs died at
around 3 hours and 8 hours with the error above.
We've been in contact with TACC and they said it was a Cactus issue, so I
am sending this report.
Attached is the parameter file I used for the tests. They should work
with a stock ET_2013_11 checkout.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1547>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#980: Remove support for the flesh-based MPI mechanism
-------------------------+--------------------------------------------------
Reporter: hinder | Owner:
Type: enhancement | Status: new
Priority: major | Milestone:
Component: Cactus | Version:
Keywords: |
-------------------------+--------------------------------------------------
The flesh-based MPI mechanism has just been deprecated in favour of
ExternalLibraries/MPI. Even though the latter is new, it is probably a
good idea to completely disable the old mechanism since having any mixture
of thorns/optionlists using the old and new mechanisms is completely
untested and will likely lead to problems and confusion. It's better to
give a useful fatal error message if someone still specifies "MPI = " in
their optionlist than to have things break in other weird and wonderful
ways.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/980>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#64: Refactory/redesign archiving
------------------------+---------------------------------------------------
Reporter: mthomas | Owner: mthomas
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
Implement archiving using archive machines with an iomethod key. Provide
another key like rsync-excludes for people to exclude files from being
archived. Provide a lightweight archive-like method for copying a
simulation from one machine to another machine.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/64>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#984: SimFactory should store the optionlist used when building in the simulation
directory
------------------------+---------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
SimFactory should store the optionlist used in the simulation directory
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/984>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#583: PITTNullCode lacks test case outputs
-----------------------------------+----------------------------------------
Reporter: eschnett | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: |
-----------------------------------+----------------------------------------
The PITTNullCode arrangement has several test parameter files without
associated output.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/583>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1390: parameter file parser aborts when findeing first error
--------------------+-------------------------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Cactus | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
The attached parfile contains multiple errors (one per line).
However the parameter file parser only reports the first one, then stops.
This makes verifying parfiles for correctness hard. It might be good to
defer aborting until the end of the file or until a larger number of
parsing errors were encountered.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1390>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1172: Remove unnecessary exp/log calls in EOS_Omni
-----------------------------------+----------------------------------------
Reporter: eschnett | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: |
-----------------------------------+----------------------------------------
EOS_Omni seems to call exp/log more often than necessary in the nuc_eos
table lookup routines. Use algebraic identities to remove them.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1172>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#614: relative tolerence in test.ccl of QuasiLocalMeasures very high
-----------------------------------+----------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: optional | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: |
-----------------------------------+----------------------------------------
LSUThorns/QuasiLocalMeasures/test/test.ccl curerntly reads:
{{{
ABSTOL 1.e-7
RELTOL 1.e+5
}}}
I am curious: is the relative tolerance of 10,000 intentional or should it
have been 1e-5 instead?
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/614>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1276: Intel 2013.1.117 mis-compiles NewRad
-----------------------------------+----------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: NewRad |
-----------------------------------+----------------------------------------
Intels compiler fails (with -O2) to push values for bmin onto the stack in
lines 316 of newrad.cc and line 126 of extrap.cc. Adding printf's for bmin
perturbs the bug out of existence, but adding a printf of the address of
bmax and reveals that at the time extrap_kernel is call the integer just
before this address is still the initialization value of bmin[2] and not
the correct value.
The attached patch disables optimization for the two driver functions
affected (but not the actual kernel).
The patch is specific (via an #if) for this particular compiler and
version. What is the best way of handling this? Target any intel version
starting from the known failing one until we know of known good one? Or
starting from an older known good one (that would be intel 11 in my case).
Hopefully no similar bug is triggered by Carpet's use of the same idiom.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1276>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1439: SSL certificate check failing
--------------------+-------------------------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Other | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
The check for SSL certificates in line 535:
{{{#perl
# check for svn SSL problems
if ( $rec{"TYPE"} eq "svn" && defined $rec{"AUTH_URL"} ) {
my $base = $rec{"AUTH_URL"};
$base =~ s/(https\:\/\/[\w\.]+)\/(.*)$/$1/i;
unless ( defined $svn_servers{$base} ) {
my $ret = `$svn --non-interactive info $rec{AUTH_URL} 2>&1`;
if ( $ret =~ /Server certificate verification failed/ ) {
$svn_servers{$base} = 0;
}
else {
$svn_servers{$base} = 1;
}
}
}
}}}
is incorrect since eg for the ET manifest where
{{{
AUTH_URL=https://svn.einsteintoolkit.org/$1/trunk
}}}
the executed svn command is:
{{{
svn --non-interactive info https://svn.einsteintoolkit.org/$1/trunk 2>&1
}}}
which actually returns and error:
{{{
svn: E175002: Unable to connect to a repository at URL
'https://svn.einsteintoolkit.org/trunk'
svn: E175002: The OPTIONS request returned invalid XML in the response:
XML parse error at line 1: Extra content at the end of the document
(https://svn.einsteintoolkit.org/trunk)
}}}
but the code does not test for svn failures at all at this point.
The simplest fix would be to move the check further down where {{{$1}}}
has been replaced by an actual value, eg into the loop:
{{{
# we are splitting each group of components into individuals
# to check for existence. they will now be passed individually to
# the checkout/update subroutines. this will take up more memory,
# but it should make it easier if the user decides to add another
# component from the same repository later
my @checkouts = split( /\s+/m, $rec{"CHECKOUT"} );
foreach my $checkout (@checkouts) {
}}}
in line 565 which however causes the test to run for every single CHECKOUT
item.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1439>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit