#1326: running loopcontrol on strange number of threads fails
-----------------------------------+----------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: LoopControl |
-----------------------------------+----------------------------------------
my machine has 8 cores (according to /proc/cpuinfo). Running eg the
trigger test with 3 threads fails inside of loopcontrol.
To reproduce:
{{{
export OMP_NUM_THREADS=3
mpirun -n 2 exe/cactus_bns_all
arrangements/AEIThorns/Trigger/test/trigger.par
}}}
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1326>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1361: disable hyperthreading in loopcontrol by default
-----------------------------------+----------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: LoopControl |
-----------------------------------+----------------------------------------
when running with both openmp and sufficiently many threads that
hyperthreading threads are used, many tests using LoopControl (Cartoon,
RotatingSymmety180, RotatingSymmetry90) fail.
This can be tracked down to disabling hyperthreading support in
LoopControl (ie. turning of hyperthreading makes things work).
In particular on bethe with smt and 8 physical cores:
The Cartoon/test_cartoon_2.par test shows differences from the recorded
results when run with 16 threads (but not with 8 threads). If I then go
ahead and disable OMP in all ML source files but ML_BSSN_enforce *and*
comment out the #include "loopcontrol.h", then the difference goes away.
Adding back #include "loopcontrol.h" brings back the error.
Some further experimenting with LoopControl's options shows that indeed
the use_smt_threads option is what causes problems. If I turn it off
things work fine even with a vanilla source tree. Otherwise relative
differences are on the order 1e-7 and absolute 1e-11 (in
momx_z_[2][2].xg). Without smt the results are identical to the stored
values.
The issue only occurs in combination of OpenMP, vectorization and
hyperthreading. The issue is independent of the compiler (both intel 13
and gcc 4.4 show the same behaviour), and vectorization (sse2) and many
threads (up to 4 times the number of physical cores) works fine on non-smt
machines.
I propose to disable LoopControl::use_smt_threads by default. Note that we
cannot completely remove it since apparently for Vesta (a Blue Gene/Q) smt
is required to get and multi-threading at all.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1361>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1204: carpet bug
-------------------------------------+--------------------------------------
Reporter: abdik@… | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone:
Component: Carpet | Version:
Keywords: |
-------------------------------------+--------------------------------------
The latest version of carpet seems to contain a bug that affects
AMR+multipatch runs. My stderr and stdout and par file are attached. Found
by Roland and Ernazar.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1204>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#997: problem in appending output after recovery
----------------------------------------+-----------------------------------
Reporter: corvino.giovanni@… | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone:
Component: Carpet | Version:
Keywords: |
----------------------------------------+-----------------------------------
I have a problem in appending output files from Carpet. I used to produce
3d HDF5 output of grid variables
and write the output in the same directory also after recovery from
checkpoint. The new output was automatically
appended to the existing one. Now I am producing h5 output also on 2D
slices but in this case the output is overwritten
so I lost the data for all but the last recovery.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/997>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1518: Parameter parser and CCTK_ParameterSet interpret leading zeros in numbers
differently
--------------------+-------------------------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: Cactus | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
the parameter parser allows things like:
{{{
thorn::param1 = 011
thorn::param2 = 012.34
}}}
in parameter files. For floating point values this is a bit unexpected but
otherwise mostly harmless. For integers the situation is a bit more
complex since in C a leading zero is used to indicate a octal number. And
(worse) while the parameter file parser converts the string "011" to the
number 11 the Cactus call CCTK_ParameterSet will convert it to 9. The
difference is ultimately the difference between calling atof (Parser) and
strtol (CCTK_ParameterSet).
To avoid confusion it would likely be good to change CCTK_ParameterSet to
behave the way the Parameter parser does. This is a change in behaviour
compared to the pre-Piraha parser, however I suspect the number of users
that actually used octal (or hexedecimal) notation in their parameter
files is small.
The change is to change {{{inval = strtol (value, &endptr, 0);}}} to
{{{inval = strtol (value, &endptr, 10);}}} in line 2209 of
src/main/Parameters.c and similar in line 2270.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1518>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#428: Forbid configuration names ending in -reconfig
----------------------+-----------------------------------------------------
Reporter: eschnett | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: Cactus | Version:
Keywords: |
----------------------+-----------------------------------------------------
I accidentally created a Cactus configuration with a name like "sim-
reconfig". This confuses Cactus, because the command "make sim-reconfig"
can then mean either to build the "sim-reconfig" configuration, or to
reconfigure the "sim" configuration.
Cactus should catch and forbid these cases.
This is probably most cleanly handled by using a script instead of a
Makefile to interpret the user commands.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/428>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1039: Web-based documentation should be automatically updated
-------------------------------------+--------------------------------------
Reporter: hinder | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit website | Version:
Keywords: Documentation |
-------------------------------------+--------------------------------------
We have documentation for Cactus and the ET on their respective websites.
This should be updated automatically when the source is changed, so it is
never out of date. There should be documentation for the current release
as well as the current development version.
There is one technical problem standing in the way of this. The system
should be automated, and it needs to have commit rights to a specific
directory in the SVN repository hosting the files. SVN accounts for CCT
machines tend to be the same as CCT login accounts, so we don't want to
store the username and password. I think the best solution is to create
an SVN account specifically for the documentation build system which is
not tied to a CCT login account. This account would be given commit
access to just the directories necessary.
The "checkout, build doc, commit" script should be kept in version control
somewhere, and a machine should be chosen on which to run it. It could be
run regularly, or just in response to commits.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1039>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1487: Documentation build should be tested as part of the automated build and
test
--------------------+-------------------------------------------------------
Reporter: hinder | Owner: hinder
Type: defect | Status: new
Priority: minor | Milestone:
Component: Other | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
The build of the ET documentation should be tested as part of the
automated build and test.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1487>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#791: Output timer tree as XML
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Carpet | Version:
Keywords: |
-------------------------+--------------------------------------------------
The attached patch to Carpet adds a parameter (off by default) which
outputs the timer tree from each process to an XML file in the output
directory at the end of the run. The output looks like this:
{{{
<timer name = "main"> 7.7702
<timer name = "CallFunction"> 0.000241
<timer name = "thorns"> 0.000233
<timer name = "CaKernel_FreeDevMem"> 0.000221
<timer name = "PostCall"> 2e-06 </timer>
<timer name = "PreCall"> 1e-06 </timer>
</timer>
</timer>
</timer>
<timer name = "CarpetStartup"> 0.00982
<timer name = "AllocateGridHierarchy"> 5e-06 </timer>
}}}
An alternative schema would be to have the name and timer value as
subelements; i.e.
{{{
<timer>
<name>main</name>
<value>7.7702</value>
<children>
<timer>
<name>CallFunction</name>
<value<0.000241</name>
...
</timer>
}}}
The <children> tag might not be necessary, but might make it easier to
parse. Having one file per process is not ideal; we would like to add
reductions across processes to give min, max, average, standard deviation
etc, also for the standard output display.
OK to commit as a work-in-progress? We can change the schema later if it
turns out to be easier to deal with.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/791>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#727: Test output changes every time some tests are run
-----------------------------------------------------------+----------------
Reporter: hinder | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: testsuites QuasiLocalMeasures IDAxiOddBrillBH |
-----------------------------------------------------------+----------------
Each time a new run of the test suites is performed, the test output of
IDAxiOddBrillBH and QuasiLocalMeasures changes a little. These changes
are below the tolerances set in the test.ccl files, so the tests pass, but
there should be no change at all when the tests are run on the same
machine.
An example of the difference from one test run to the next is shown at
http://git.barrywardell.net/EinsteinToolkitTestResults.git/blobdiff/95fe088….
Another example, from QuasiLocalMeasures, is at
http://git.barrywardell.net/EinsteinToolkitTestResults.git/blobdiff/95fe088…
/qlm-ks-shifted/admbase::metric.average.asc.
The only thorns to exhibit this behaviour are IDAxiOddBrillBH and
QuasiLocalMeasures. All other test data remains unchanged from one test
run to the next. Note that the two test runs will be performed with
different executables, since a new configuration is created for each run.
I think I remember testing that simply re-running the same executable
resulted in the same test output.
Is it possible that there are a small number of points which access
uninitialised memory which changes on each run for these thorns?
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/727>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit