#1373: GRHydro::tov_slowsector test fails in ET_2013_05
-----------------------------------+----------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: GRHydro |
-----------------------------------+----------------------------------------
Most likely due to intel vs. gcc compiler issues. This ticket to serve as
a reminder to fix this and to collect information known about this issue.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1373>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1378: provide equivalent of fprintf(stderr, "%s\n", msg) in Fortran
-------------------------+--------------------------------------------------
Reporter: rhaas | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Cactus | Version: development version
Keywords: |
-------------------------+--------------------------------------------------
Currently to write out multi-line error messages in Fortran we use
multiple calls to CCTK_WARN(1, warnline) followed possibly by a
CCTK_ERROR(errline). Each of the level 1 warnings (given certain settings
of the Cactus parameters) prints the source file location and other
information to screen, thus cluttering the error output. A typical error
message might look like this:
{{{
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 386 of GRHydro_Prim2Con.F90):
-> EOS error in prim2con_hot:
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 388 of GRHydro_Prim2Con.F90):
-> 64897 22 37 31 -1.440000E+00 -8.496000E+01 -1.296000E+01
8.595485E+01
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 390 of GRHydro_Prim2Con.F90):
-> 1.228064E-09 -8.644951E-03 -5.842300E-01 4.789413E-01
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 392 of GRHydro_Prim2Con.F90):
-> code: 106
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 394 of GRHydro_Prim2Con.F90):
-> reflevel: 0
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 386 of GRHydro_Prim2Con.F90):
-> EOS error in prim2con_hot:
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 388 of GRHydro_Prim2Con.F90):
-> 64897 23 37 31 1.440000E+00 -8.496000E+01 -1.296000E+01
8.595485E+01
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 390 of GRHydro_Prim2Con.F90):
-> 1.227982E-09 -8.644755E-03 -5.798389E-01 4.789407E-01
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 392 of GRHydro_Prim2Con.F90):
-> code: 106
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 394 of GRHydro_Prim2Con.F90):
-> reflevel: 0
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 166 of GRHydro_Eigenproblem.F90):
-> EOS ERROR in eigenvalues_hot
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 168 of GRHydro_Eigenproblem.F90):
-> keyerr: 668 keytemp: 0
WARNING level 0 in thorn GRHydro processor 179 host nid03278
(line 170 of GRHydro_Eigenproblem.F90):
-> 1.228064E-09 -8.644951E-03 -5.842300E-01 4.789413E-01
5.668696E-04
cactus_sim:
Cactus/arrangements/Carpet/Carpet/src/helpers.cc:314:
int Carpet::Abort(const cGH*, int): Assertion `0' failed.
Rank 179 with PID 13388 received signal 6
}}}
with errors from multiple MPI processes possibly intersecting each other.
It would be useful to provide a subroutine equivalent to
{{{
subroutine CCTK_WARN_SHORT(msg)
character*(*) :: msg
write (stderr,'(a)') msg
end subroutine
}}}
(name is up for discussion) that outputs only "msg" to stderr (and the
warning listener registered in the flesh) without prepending the file
information output.
A similar routine might be offered for C (in WARN and VWarn flavors) both
for symmetry reasons and to have the message pass the warning listeners,
though in C one can usually get away with a single CCTK_VWarn and a very
long format string.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1378>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1551: Strange warnings at startup
----------------------+-----------------------------------------------------
Reporter: eschnett | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: Cactus | Version: development version
Keywords: |
----------------------+-----------------------------------------------------
I see this warning multiple times at startup on Blue Waters:
{{{
^[[1mWARNING[L1,P0] (Flesh):^[[0m Invalid end
}}}
This is apparently output by the routine Util_IntInRange or
Util_DoubleInRange. This happens for the parameter file
simfactory/etc/parfiles/submit.par.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1551>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1473: Carpet segfaults in Shutdown
--------------------+-------------------------------------------------------
Reporter: knarf | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: Carpet | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
For high numbers of threads, Carpet segfaults in Shutdown for the
testsuite CarpetWaveToyNewRecover_test_1proc. I tried this with gcc and
intel on a Debian system. The machine I run this on (spine, with the
default simfactory configurations) has 2 processors with 8 cores, and ht
enabled. When I run this testsuite with 1 mpi process but certain numbers
of threads (e.g. 9, 11, 16) I see this segfault. Which number triggers the
bug seems to depend on the compiler, but seems to be consistent within
tries.
The backtrace I see is, e.g.,:
{{{
1. CarpetLib::signal_handler(int)
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN9CarpetLib14signal_handlerEi+0xeb)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/backtrace.cc:542]
2. /lib/x86_64-linux-gnu/libc.so.6(+0x324f0) [??:0]
3. /lib/x86_64-linux-gnu/libc.so.6(gsignal+0x35) [??:0]
4. /lib/x86_64-linux-gnu/libc.so.6(abort+0x180) [??:0]
5. /lib/x86_64-linux-gnu/libc.so.6(+0x6d52b) [??:0]
6. /lib/x86_64-linux-gnu/libc.so.6(+0x76d76) [??:0]
7. /lib/x86_64-linux-gnu/libc.so.6(cfree+0x6c) [??:0]
8. mem<double>::~mem()
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN3memIdED1Ev+0x80)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/mem.cc:185]
9. data<double>::free()
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN4dataIdE4freeEv+0x4b)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/data.cc:549]
a. data<double>::~data()
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN4dataIdED1Ev+0x27)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/data.cc:487]
b. data<double>::~data()
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN4dataIdED0Ev+0x9)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/data.cc:487]
c. ggf::recompose_free(int)
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN3ggf14recompose_freeEi+0x1ae)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/ggf.cc:257]
d. ggf::~ggf() [/home/knarf/ET_2013_11/exe/cactus_sim(_ZN3ggfD1Ev+0xde)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/ggf.cc:70]
e. gf<double>::~gf()
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN2gfIdED0Ev+0x9)
[/home/knarf/ET_2013_11/configs/sim/build/CarpetLib/gf.cc:30]
f. Carpet::Shutdown(tFleshConfig*)
[/home/knarf/ET_2013_11/exe/cactus_sim(_ZN6Carpet8ShutdownEP12tFleshConfig+0x3ad)
[/home/knarf/ET_2013_11/configs/sim/build/Carpet/Shutdown.cc:102]
10. /home/knarf/ET_2013_11/exe/cactus_sim(main+0x49)
[/home/knarf/ET_2013_11/configs/sim/build/Cactus/main/flesh.cc:92]
11. /lib/x86_64-linux-gnu/libc.so.6(__libc_start_main+0xfd) [??:0]
12. /home/knarf/ET_2013_11/exe/cactus_sim() [??:0]
}}}
Marking as minor because this is in Shutdown, however if something would
actually check for this it might break workflows.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1473>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1549: jobs starting/ending on stampede do not send emails
------------------------+---------------------------------------------------
Reporter: anonymous | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version: development version
Keywords: |
------------------------+---------------------------------------------------
jobs starting/ending on stampede do not send emails because stampede.sub
lacks a line like:
#SBATCH --mail-user=
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1549>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1541: WaveToy2D outputs what looks like unitiialized data
--------------------------+-------------------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: Cactus | Version: development version
Keywords: WaveToy2DF77 |
--------------------------+-------------------------------------------------
the test test_WaveToy2D in WaveToy2DF77 in the CactusExamples arrangement
fails for me (with a [mostly] clean checkout):
{{{
WaveToy2DF77: test_WaveToy2D
Failure: 7 files compared, 6 differ, 6 differ significantly
}}}
attached are diff log and one of the data files. This happens with current
trunk.
It looks to me as if I am seeing poison.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1541>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#590: McLachlan should allow other thorns to set the gauge
-----------------------------------+----------------------------------------
Reporter: bmundim | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: |
-----------------------------------+----------------------------------------
The parameters lapse_evolution_method and shift_evolution_method are
usually set to ML_BSSN in McLachlan. However they are never checked
in the code. McLachlan indeed seems to ignore their values and
overwrite whatever the values of lapse or shift set elsewhere,
preventing therefore other thorns from setting them differently.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/590>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1495: Hopper doesn't pass test suite
------------------------+---------------------------------------------------
Reporter: eschnett | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version: development version
Keywords: |
------------------------+---------------------------------------------------
There are many test suite failures on Hopper.
The problem seems to be MPI errors when running on multiple nodes. I do
not understand what is going wrong. I suspect a problem with our build or
run options. I have asked NERSC for help.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1495>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1497: CarpetInterp/waveinterp_1p and CarpetInterp/waveinterp_2p tests are failing
--------------------+-------------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: Carpet | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
The old tests CarpetInterp/waveinterp_1p and CarpetInterp/waveinterp_2p,
which are almost identical apart from running on 1 or 2 processes, are
both failing with the error
{{{
----------------------------------------------------------------
Iteration Time | WAVEMOL::phi | *VEMOL::phit | *VEMOL::phix
| norm2 | norm2 | norm2
----------------------------------------------------------------
0 0.138 | -nan | -nan | -nan
WARNING level 0 in thorn CarpetLib processor 0 host
caltechcactusjenkins.spdns.org
(line 428 of
/home/jenkins/workspace/EinsteinToolkit/arrangements/Carpet/CarpetLib/src/gdata.cc):
-> Internal error: extrapolation in time. variable=WAVEMOL::phi
time=0.20000000000000001
times=[0.40000000000000002,0.67500000000000004,0.27500000000000002]
WARNING level 0 in thorn CarpetLib processor 0 host
caltechcactusjenkins.spdns.org
(line 428 of
/home/jenkins/workspace/EinsteinToolkit/arrangements/Carpet/CarpetLib/src/gdata.cc):
-> Internal error: extrapolation in time. variable=WAVEMOL::phi
time=0.20000000000000001
times=[0.40000000000000002,0.67500000000000004,0.27500000000000002]
}}}
This test is newly failing because the dependent thorn WaveMoL has only
just been added to the toolkit. The tests had a parameter setting
WaveMoL::num_timelevels = 3 which didn't correspond to any WaveMoL
parameter, so previously the tests would crash earlier. I removed that
parameter setting, and now the tests fail with this "extrapolation in
time" error. I don't understand this failure. The times array seems to
be non-monotonic [0.4, 0.675, 0.275] perhaps indicating that it might not
have been initialised correctly? The coarse grid timestep should be
0.1*0.25 = 0.025, so I don't see where these times come from during
iteration 1. The NaNs in the norms are probably poison, which probably
wasn't enabled when this test was written, so nobody noticed.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1497>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1340: simfactory does not abort --testsuite submission process if rsync fails
------------------------+---------------------------------------------------
Reporter: rhaas | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
when setting up testsuite runs simfactory uses rsync to copy the test
suite data into the simulation folder. If this rsync fails (eg. because a
user specified incorrect rsyncopts in defs.local.ini) the submission
process does not abort and instead submits an emtpy test-suite run.
{{{
rhaas@kraken-gsi2:~/ET_trunk> sim create-submit 2p6t --procs 12 --num-
threads 6 --walltime 4:0:0 --tests
uite --allocation TG-ASC120003
Skeleton Created
Job directory: "/lustre/scratch/rhaas/simulations/2p6t"
Option --testsuite given
Executable: "/nics/c/home/rhaas/ET_trunk/exe/cactus_sim"
Option list:
"/lustre/scratch/rhaas/simulations/2p6t/SIMFACTORY/cfg/OptionList"
Submit script:
"/lustre/scratch/rhaas/simulations/2p6t/SIMFACTORY/run/SubmitScript"
Run script:
"/lustre/scratch/rhaas/simulations/2p6t/SIMFACTORY/run/RunScript"
Assigned restart id: 0
Copying testsuite data
rsync: --times=no: option does not take an argument
rsync error: syntax or usage error (code 1) at main.c(1435) [client=3.0.9]
Executing submit command: /opt/torque/2.5.7/bin/qsub
/lustre/scratch/rhaas/simulations/2p6t/output-0000/SIMFACTORY/SubmitScript
Submit finished, job id is 3236567.nid00016
rhaas@kraken-gsi2:~/ET_trunk> qdel 3236567.nid00016
}}}
My rsynopts were:
{{{
rsyncopts = --times=no --checksum --include 'configs/*/ThornList'
--exclude 'configs/*/*'
}}}
which are bad for two reasons:
1.) kraken's rsync does not no --times-no (likely wants --notimes or so)
2.) --exclude 'configs/*/*' excludes cctk_MPI.h which is used by the test
suite infrastructure to detect the presence of MPI
Note that some of these options are obviously obsolete now that simfactory
defaults to --times=no --checksum anyway.
Still, simfactory should always check the exit status of any command it
calls I think.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1340>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit