#1473: Carpet segfaults in Shutdown
--------------------+-------------------------------------------------------
Reporter: knarf | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: Carpet | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
For high numbers of threads, Carpet segfaults in Shutdown for the
testsuite CarpetWaveToyNewRecover_test_1proc. I tried this with gcc and
intel on a Debian system. The machine I run this on (spine, with the
default simfactory configurations) has 2 processors with 8 cores, and ht
enabled. When I run this testsuite with 1 mpi process but certain numbers
of threads (e.g. 9, 11, 16) I see this segfault. Which number triggers the
bug seems to depend on the compiler, but seems to be consistent within
tries.
The backtrace I see is, e.g.,:
{{{
1. CarpetLib::signal_handler(int)
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN9CarpetLib14signal_handlerEi+0xeb)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/backtrace.cc:542]
2. /lib/x86_64-linux-gnu/libc.so.6(+0x324f0) [??:0]
3. /lib/x86_64-linux-gnu/libc.so.6(gsignal+0x35) [??:0]
4. /lib/x86_64-linux-gnu/libc.so.6(abort+0x180) [??:0]
5. /lib/x86_64-linux-gnu/libc.so.6(+0x6d52b) [??:0]
6. /lib/x86_64-linux-gnu/libc.so.6(+0x76d76) [??:0]
7. /lib/x86_64-linux-gnu/libc.so.6(cfree+0x6c) [??:0]
8. mem<double>::~mem()
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN3memIdED1Ev+0x80)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/mem.cc:185]
9. data<double>::free()
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN4dataIdE4freeEv+0x4b)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/data.cc:549]
a. data<double>::~data()
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN4dataIdED1Ev+0x27)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/data.cc:487]
b. data<double>::~data()
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN4dataIdED0Ev+0x9)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/data.cc:487]
c. ggf::recompose_free(int)
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN3ggf14recompose_freeEi+0x1ae)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/ggf.cc:257]
d. ggf::~ggf() [/home/knarf/ET_2013_11/exe/cactus_sim(_ZN3ggfD1Ev+0xde)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/ggf.cc:70]
e. gf<double>::~gf()
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN2gfIdED0Ev+0x9)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/gf.cc:30]
f. Carpet::Shutdown(tFleshConfig*)
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN6Carpet8ShutdownEP12tFleshConfig+0x3ad)
[/home/knarf/ET_2013_11/configs/sim/build/Carpet/Shutdown.cc:102]
10. /home/knarf/ET_2013_11/exe/cactus_sim(main+0x49)
[/home/knarf/ET_2013_11/configs/sim/build/Cactus/main/flesh.cc:92]
11. /lib/x86_64-linux-gnu/libc.so.6(__libc_start_main+0xfd) [??:0]
12. /home/knarf/ET_2013_11/exe/cactus_sim() [??:0]
}}}
Marking as minor because this is in Shutdown, however if something would
actually check for this it might break workflows.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1473>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1576: use https as URL for repositories where available
-------------------------+--------------------------------------------------
Reporter: knarf | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone: ET_2014_11
Component: Other | Version: development version
Keywords: |
-------------------------+--------------------------------------------------
It would be nice to not have different access URLs for repositories
depending on whether a user has write access or not. This too often leads
to problems when checking something out as "read only" and trying to
commit/push later.
I propose to have only one URL for repositories where this is possible.
For most svn repositories this would mean not using plain http, and for
git it would mean combining git@ and git:// into https:// (where possible,
carpetcode-hosted repositories do not provide this right now).
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1576>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1604: Uninitialised data in tests using PUGH
-----------------------------------+----------------------------------------
Reporter: hinder | Owner:
Type: defect | Status: new
Priority: major | Milestone: ET_2014_05
Component: EinsteinToolkit thorn | Version: development version
Keywords: |
-----------------------------------+----------------------------------------
PUGH has a poisoning parameter which can initialise grid variables to NaN.
The default is "none" which leaves grid variables uninitialised. If the
default is changed to NaN, several tests fail, indicating that they are
failing to initialise their data correctly. The tests which fail are:
tov_slowsector (from GRHydro)
TestComplex (from TestComplex)
gaussian (from WaveMoL)
The tests fail on both 1 and 2 processes.
All report differences related to NaNs apart from tov_slowsector.
Possibly there is some NaN-filtering going on there. The test passes
without PUGH poisoning. Tests were performed on Datura.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1604>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1602: captcha not working anymore
----------------------------------+-----------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: EinsteinToolkit trac | Version: development version
Keywords: Captcha |
----------------------------------+-----------------------------------------
it seems as if the Captcha that gets triggered by potential spam is no
longer triggered and the user is just presented by a page that claims his
comment is spam.
This just happened to Bela Szilagyi whom we had asked for input on ticket
#1600. Please if someone at LSU would take a look at this this would be
very helpful.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1602>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1583: Carpet default poison value should not be NaN
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: enhancement | Status: new
Priority: optional | Milestone:
Component: Carpet | Version: development version
Keywords: |
-------------------------+--------------------------------------------------
Carpet initialises Cactus variables to NaN as this is a quantity which is
likely to be noticed. However, this makes it impossible to know when you
see a NaN whether you have a programming error (accessing uninitialised
memory) or a numerical problem (numerical solution has blown up).
I usually set the carpet poison value to something else; a value like
10^230 is just as likely to be noticed as a NaN, and when you see exactly
that value, you know that you are looking at uninitialised data, rather
than a computation which went wrong.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1583>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1603: CarpetProlongateTest/test_o9 is failing on several machines
--------------------+-------------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone: ET_2014_05
Component: Carpet | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
CarpetProlongateTest/test_o9 is failing on Gordon, Mike and Redshift. On
Gordon, the differences are
{{{
carpetprolongatetest::difference.d.asc: substantial differences
significant differences on 1 (out of 206) lines
maximum absolute difference in column 13 is 1.07288360595703e-06
maximum absolute difference in column 14 is 6.98491930961609e-10
maximum relative difference in column 13 is 1.12625729345722e-12
maximum relative difference in column 14 is 3
(insignificant differences on 19 lines)
carpetprolongatetest::difference.x.asc: differences below tolerance on
26 lines
carpetprolongatetest::difference.y.asc: differences below tolerance on
26 lines
carpetprolongatetest::errornorm..asc: differences below tolerance on 1
lines
carpetprolongatetest::scalar.d.asc: differences below tolerance on 13
lines
carpetprolongatetest::scalar.x.asc: differences below tolerance on 6
lines
carpetprolongatetest::scalar.y.asc: differences below tolerance on 4
lines
}}}
and on the other machines the results are similar. test.ccl contains
{{{
TEST test_o7
{
ABSTOL 2.0e-11
}
TEST test_o9
{
ABSTOL 5.0e-10
}
TEST test_o11
{
ABSTOL 3.0e-8
}
}}}
Higher order prolongation probably leads to more amplification of roundoff
differences, which is why the absolute tolerances listed here increase
with prolongation order.
On Redshift, which uses -Ofast with gcc, the maximum absolute differences
in columns 13 and 14 are just marginally above the tolerance of 5e-10:
{{{
maximum absolute difference in column 13 is 9.31322574615479e-10
maximum absolute difference in column 14 is 6.98491930961609e-10
maximum relative difference in column 13 is 9.77653900570505e-16
maximum relative difference in column 14 is 3
}}}
However on Gordon and Mike, the column 13 absolute difference is 1e-6,
which presumably means the data is large, so the relative tolerance will
come into play. The default relative tolerance is 1e-12, and the
difference in column 13 is marginally above this.
Gordon:
{{{
maximum absolute difference in column 13 is 1.07288360595703e-06
maximum absolute difference in column 14 is 6.98491930961609e-10
maximum relative difference in column 13 is 1.12625729345722e-12
maximum relative difference in column 14 is 3
}}}
Mike:
{{{
maximum absolute difference in column 13 is 1.07288360595703e-06
maximum absolute difference in column 14 is 6.98491930961609e-10
maximum relative difference in column 13 is 1.12625729345722e-12
maximum relative difference in column 14 is 3
}}}
Should we increase both the relative and absolute tolerances for this test
to 1e-11? I believe that would make the test pass on all three machines.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1603>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1591: Simfactory: add -Wall and '-warn all' options to warning flags for intel
compilers
------------------------+---------------------------------------------------
Reporter: bmundim | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone: ET_2014_05
Component: SimFactory | Version: development version
Keywords: |
------------------------+---------------------------------------------------
If you add the flag '-warn all' to F77_WARN_FLAGS and F90_WARN_FLAGS the
compiler catches a subtle error (compiler error #8284) related to
subroutine and function calls in Fortran:
{{{
https://software.intel.com/en-us/forums/topic/342615https://software.intel.com/en-us/blogs/2009/03/31/doctor-fortran-in-ive-
come-here-for-an-argument
}}}
Currently grhydro won't compile if we turn on that flag. I suggest to
update all configurations in simfactory using intel compilers with the
flag mentioned above and fix any issue related to this bug on the ET
thorns. Any thoughts on my
proposal?
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1591>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1601: test system ignores nprocs unless MPI is found
--------------------+-------------------------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: Cactus | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
the test system script RunTestUtils.pl checks if the configuration is MPI
enabled by looking into some build files (see function ParseExtras and its
use of file $config/bindings/Configuration/Capabilities/cctki_MPI.h).
However in line 549
{{{
if ($have_mpi)
{
$config_data->{'NPROCS'} =
&defprompt(' Enter number of processors ($nprocs)', $nprocs);
}
else
{
print "No MPI available\n";
if ($nprocs > 1)
{
die("Cannot run on $nprocs processes without an MPI
implementation\n");
}
}
}}}
it only sets $config_data->{NPROCS} if MPI was found and leaves it
undefined otherwise. This seems to confuse later parts of the script which
try to decide which tests to run and one gets eg
{{{
Summary for configuration sim
Time -> Mon Apr 28 21:35:38 PDT 2014
Host -> nid27637
Processes ->
User -> rhaas
Total available tests -> 287
Unrunnable tests -> 133
Runnable tests -> 154
Total number of thorns -> 209
Number of tested thorns -> 52
Number of tests passed -> 149
Number passed only to
set tolerance -> 95
Number failed -> 5
}}}
and
{{{
Tests missed for different number of processors required:
checkpointML-EE in AHFinderDirect
(EinsteinAnalysis/AHFinderDirect/test/checkpointML-EE.par)
Requires 1 processors
test_cc_o5 in CarpetProlongateTest
(CarpetExtra/CarpetProlongateTest/test/test_cc_o5.par)
Requires 2 processors
}}}
I suggest setting congif_data->{NPROCS} to 1 if no MPI is found.
Note that this really only happens if one manages to compile without MPI
(which implies no Carpet) or (as in my case) uses simfactory (with 1
process) but has rsync filter rules that prevent it from copying all of
configs/sim to the simulation base dir.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1601>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1605: Multipole convergence order test is sensitive to roundoff
-----------------------------------+----------------------------------------
Reporter: hinder | Owner:
Type: defect | Status: new
Priority: major | Milestone: ET_2014_05
Component: EinsteinToolkit thorn | Version: development version
Keywords: |
-----------------------------------+----------------------------------------
The integration convergence test performed by Multipole is sensitive to
roundoff, as the relative error is ~1e-7 in the Simpson case. This is
important only during the correctness test, not for general use of the
thorn. The computation of the convergence order includes a subtraction of
the exact from the numerical result, and this loses many digits of
precision in the final convergence order. This causes the simpson test to
fail on several machines. So, while it is useful to have the convergence
order output, this is not a good regression test. The attached set of
patches output the integration results, and raise the tolerance of the
convergence order to 1e-3. The tolerance for the integration results is
left at the default. The tests all pass on Datura, though they also did
before this change.
OK to commit?
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1605>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1549: jobs starting/ending on stampede do not send emails
------------------------+---------------------------------------------------
Reporter: anonymous | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version: development version
Keywords: |
------------------------+---------------------------------------------------
jobs starting/ending on stampede do not send emails because stampede.sub
lacks a line like:
#SBATCH --mail-user=
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1549>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit