#499: Prolongation fails with vectorisation enabled
--------------------+-------------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone:
Component: Carpet | Version:
Keywords: |
--------------------+-------------------------------------------------------
The development version of Carpet uses vectorisation to speed-up
prolongation. This fails with various errors, including corruption of the
malloc heap.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/499>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#445: Make Carpet timers hierarchical
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: enhancement | Status: new
Priority: major | Milestone:
Component: Carpet | Version:
Keywords: |
-------------------------+--------------------------------------------------
With the current flat structure of timers in Carpet, it is difficult to
identify which timers are contained in which other timers, and hence to
avoid double-counting when adding up the times.
This series of patches modifies the timer infrastructure in Carpet to
generate a tree of timers where the hierarchy reflects the call-graph of
the program. This makes it much easier to interpret the timer output than
with the previous flat structure, where it was not possible to see which
timers "contained" which others. More implementation details are given at
the top of TimerNode.hh.
Note that the Timer source and header files have been renamed as
CactusTimer and a new Timer file and object has been created. This is
because the Timer object now only provides a wrapper around the Cactus
timer mechanism which was contained in the old Timer object.
New parameters output_initialise_timer_tree and output_timer_tree_every
control output of a new "timer tree diagram" to standard output for the
Initialise and Evolve timer trees respectively. These diagrams indicate:
1. the value of each timer;
2. the percentage of the given tree taken by each timer;
3. which timers are contained in which other timers;
4. any untimed code
for any timer which takes more than 1% of the tree time.
Making the timers hierarchical means that the ad-hoc methods used before
to identify the hierarchy (such as naming the timer Evolve::Sync, for
example, to indicate that the Sync timer was a child of the Evolve timer)
are no longer necessary and have been removed. Additionally, the
construction of timers in a "dynamic" manner is now handled automatically
for all timers, so special-case code is no longer needed and has been
removed.
Additionally some previously-untimed parts of the code are now timed, and
timer names have been made more consistent in some places.
There is code in the patches to output the entire timer tree as an XML
file, but it is not enabled.
Ideally the timer tree printed to standard output would contain reductions
across processes, but at the moment it contains only the timer from the
current process.
The attached "tree-example.txt" file shows an example of the timer tree
that is printed for a simulation using the qc0-mclachlan parameter file
from the Einstein Toolkit.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/445>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1210: CarpetIOASCII should not write column names info each iteration
--------------------+-------------------------------------------------------
Reporter: knarf | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: Carpet | Version:
Keywords: |
--------------------+-------------------------------------------------------
Currently, CarpetIOASCII writes the column names header each iteration,
leading to bloated output - easily double the size of what should suffice.
Once after file creation would be much better.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1210>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1399: add evolution method "stationary" to ADMBase
-----------------------------------+----------------------------------------
Reporter: rhaas | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version: development version
Keywords: ADMBase |
-----------------------------------+----------------------------------------
This method is similar to "static" in that it can be used for Cowling
runs, it differs from "static" in that it schedules the initial data
routine after any grid changes and does not apply boundary conditions or
SYNC calls. It can thus be used if one can compute the required
metric/shift/lapse easily in a pointwise manner and would like to have
exact (non-prolongated) data in the buffer zones.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1399>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1302: Reduce time spent in deciding not to do output.
--------------------------------------+-------------------------------------
Reporter: diener | Owner: eschnett
Type: enhancement | Status: new
Priority: major | Milestone:
Component: Carpet | Version: development version
Keywords: performance optimization |
--------------------------------------+-------------------------------------
I noticed, during some scaling tests, that for my code a significant
amount of time was spent by CarpetIOASCII, even though I only had it
activated and didn't actually request any output. The reason is that in
order to figure out whether to do output or not, there is a loop over all
grid variables and a routine (TimeToOutput) is called. In this
routine there is a check if the out_dir and out_vars parameters have been
steered and if so update some internal data structures. It should be
sufficient to do this before entering the loop over grid variables. The
same issue is present in CarpetIOScalar and CarpetIOHDF5. The attached
patches for CarpetIOASCII, CarpetIOScalar and CarpetIOHDF5 moves this
check outside of the loop over grid variables and in addition bypasses the
loop completely if the out_vars parameter string is the empty string. Note
that the code where I noticed this problem uses large vectors of 1D grid
arrays, and it turns out that Cactus counts each vector element as a
distinct grid variables and the length of the loop over grid variables in
my case was close to 100.000, which might explain why nobody has noticed
this before.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1302>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1373: GRHydro::tov_slowsector test fails in ET_2013_05
-----------------------------------+----------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: GRHydro |
-----------------------------------+----------------------------------------
Most likely due to intel vs. gcc compiler issues. This ticket to serve as
a reminder to fix this and to collect information known about this issue.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1373>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1378: provide equivalent of fprintf(stderr, "%s\n", msg) in Fortran
-------------------------+--------------------------------------------------
Reporter: rhaas | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Cactus | Version: development version
Keywords: |
-------------------------+--------------------------------------------------
Currently to write out multi-line error messages in Fortran we use
multiple calls to CCTK_WARN(1, warnline) followed possibly by a
CCTK_ERROR(errline). Each of the level 1 warnings (given certain settings
of the Cactus parameters) prints the source file location and other
information to screen, thus cluttering the error output. A typical error
message might look like this:
{{{
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 386 of GRHydro_Prim2Con.F90):
-> EOS error in prim2con_hot:
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 388 of GRHydro_Prim2Con.F90):
-> 64897 22 37 31 -1.440000E+00 -8.496000E+01 -1.296000E+01
8.595485E+01
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 390 of GRHydro_Prim2Con.F90):
-> 1.228064E-09 -8.644951E-03 -5.842300E-01 4.789413E-01
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 392 of GRHydro_Prim2Con.F90):
-> code: 106
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 394 of GRHydro_Prim2Con.F90):
-> reflevel: 0
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 386 of GRHydro_Prim2Con.F90):
-> EOS error in prim2con_hot:
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 388 of GRHydro_Prim2Con.F90):
-> 64897 23 37 31 1.440000E+00 -8.496000E+01 -1.296000E+01
8.595485E+01
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 390 of GRHydro_Prim2Con.F90):
-> 1.227982E-09 -8.644755E-03 -5.798389E-01 4.789407E-01
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 392 of GRHydro_Prim2Con.F90):
-> code: 106
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 394 of GRHydro_Prim2Con.F90):
-> reflevel: 0
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 166 of GRHydro_Eigenproblem.F90):
-> EOS ERROR in eigenvalues_hot
WARNING level 1 in thorn GRHydro processor 179 host nid03278
(line 168 of GRHydro_Eigenproblem.F90):
-> keyerr: 668 keytemp: 0
WARNING level 0 in thorn GRHydro processor 179 host nid03278
(line 170 of GRHydro_Eigenproblem.F90):
-> 1.228064E-09 -8.644951E-03 -5.842300E-01 4.789413E-01
5.668696E-04
cactus_sim:
Cactus/arrangements/Carpet/Carpet/src/helpers.cc:314:
int Carpet::Abort(const cGH*, int): Assertion `0' failed.
Rank 179 with PID 13388 received signal 6
}}}
with errors from multiple MPI processes possibly intersecting each other.
It would be useful to provide a subroutine equivalent to
{{{
subroutine CCTK_WARN_SHORT(msg)
character*(*) :: msg
write (stderr,'(a)') msg
end subroutine
}}}
(name is up for discussion) that outputs only "msg" to stderr (and the
warning listener registered in the flesh) without prepending the file
information output.
A similar routine might be offered for C (in WARN and VWarn flavors) both
for symmetry reasons and to have the message pass the warning listeners,
though in C one can usually get away with a single CCTK_VWarn and a very
long format string.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1378>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#590: McLachlan should allow other thorns to set the gauge
-----------------------------------+----------------------------------------
Reporter: bmundim | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: |
-----------------------------------+----------------------------------------
The parameters lapse_evolution_method and shift_evolution_method are
usually set to ML_BSSN in McLachlan. However they are never checked
in the code. McLachlan indeed seems to ignore their values and
overwrite whatever the values of lapse or shift set elsewhere,
preventing therefore other thorns from setting them differently.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/590>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1340: simfactory does not abort --testsuite submission process if rsync fails
------------------------+---------------------------------------------------
Reporter: rhaas | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
when setting up testsuite runs simfactory uses rsync to copy the test
suite data into the simulation folder. If this rsync fails (eg. because a
user specified incorrect rsyncopts in defs.local.ini) the submission
process does not abort and instead submits an emtpy test-suite run.
{{{
rhaas@kraken-gsi2:~/ET_trunk> sim create-submit 2p6t --procs 12 --num-
threads 6 --walltime 4:0:0 --tests
uite --allocation TG-ASC120003
Skeleton Created
Job directory: "/lustre/scratch/rhaas/simulations/2p6t"
Option --testsuite given
Executable: "/nics/c/home/rhaas/ET_trunk/exe/cactus_sim"
Option list:
"/lustre/scratch/rhaas/simulations/2p6t/SIMFACTORY/cfg/OptionList"
Submit script:
"/lustre/scratch/rhaas/simulations/2p6t/SIMFACTORY/run/SubmitScript"
Run script:
"/lustre/scratch/rhaas/simulations/2p6t/SIMFACTORY/run/RunScript"
Assigned restart id: 0
Copying testsuite data
rsync: --times=no: option does not take an argument
rsync error: syntax or usage error (code 1) at main.c(1435) [client=3.0.9]
Executing submit command: /opt/torque/2.5.7/bin/qsub
/lustre/scratch/rhaas/simulations/2p6t/output-0000/SIMFACTORY/SubmitScript
Submit finished, job id is 3236567.nid00016
rhaas@kraken-gsi2:~/ET_trunk> qdel 3236567.nid00016
}}}
My rsynopts were:
{{{
rsyncopts = --times=no --checksum --include 'configs/*/ThornList'
--exclude 'configs/*/*'
}}}
which are bad for two reasons:
1.) kraken's rsync does not no --times-no (likely wants --notimes or so)
2.) --exclude 'configs/*/*' excludes cctk_MPI.h which is used by the test
suite infrastructure to detect the presence of MPI
Note that some of these options are obviously obsolete now that simfactory
defaults to --times=no --checksum anyway.
Still, simfactory should always check the exit status of any command it
calls I think.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1340>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit