#2219: fix sqrt() for AVX512, updates to Vectors thorn
-----------------------------------+---------------------------------
Reporter: Roland Haas | Owner: Roland Haas
Type: defect | Status: assigned
Priority: critical | Milestone: ET_2019_02
Component: EinsteinToolkit thorn | Version: development version
Keywords: Vectors |
-----------------------------------+---------------------------------
This pull request mostly add better support for the AVX512 instructions
found in KNL and modern Intel Xeon / Core iN CPUs.
There is one bugfix commit “Correct error in AVX512 sqrt implemention”
which is important for those architectures (eg to compute sqrt(detg)).
I am marking this as critical until I can verify that without this fix the
result of a BBH simulation on eg Stampede2 is not incorrect because of
using an incorrect sqrt() function.
Pull request is here: https://bitbucket.org/cactuscode/cactusutils/pull-
requests/17/rhaas-avx512/diff
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/2219>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#2218: Simplify including external files in Formaline
---------------------------------+-----------------------------------
Reporter: Erik Schnetter | Type: enhancement
Status: new | Priority: optional
Milestone: | Component: EinsteinToolkit thorn
Version: development version | Keywords:
---------------------------------+-----------------------------------
Formaline uses a complicated mechanism to include tarballs of the thorns
into the executable, requiring running perl scripts and adding additional
source files. This can be simplified. See
<https://hsyl20.fr/home/posts/2019-01-15-fast-file-embedding-with-
ghc.html> for how to implement this.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/2218>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#2216: Vectors: allow seamless interaction with scalar types
---------------------------------+-----------------------------------
Reporter: Roland Haas | Type: enhancement
Status: new | Priority: minor
Milestone: | Component: EinsteinToolkit thorn
Version: development version | Keywords: Vectors
---------------------------------+-----------------------------------
* Vectors: allow implicit cast from scalar_t
* Vectors: provide binary operations with scalars
one needs both
{{{
operatorX(T, vectype<T>)
}}}
and
{{{
operatorX(vectype<T>, T)
}}}
for both commuting and non-commuting operators
Pull request is here: https://bitbucket.org/cactuscode/cactusutils/pull-
requests/16/rhaas-inclict-conv/diff
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/2216>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#2211: gh::regrid contains code whose runtime is quadratic in number of components
---------------------------------+--------------------
Reporter: Roland Haas | Type: defect
Status: new | Priority: minor
Milestone: | Component: Carpet
Version: development version | Keywords:
---------------------------------+--------------------
The member function gh::regrid contains this code
{{{
// Check component consistency
for (int ml = 0; ml < mglevels(); ++ml) {
for (int rl = 0; rl < reflevels(); ++rl) {
assert(components(rl) >= 0);
for (int c = 0; c < components(rl); ++c) {
ibbox const &b = extent(ml, rl, c);
ibbox const &b0 = extent(ml, rl, 0);
assert(all(b.stride() == b0.stride()));
assert(b.is_aligned_with(b0));
for (int cc = c + 1; cc < components(rl); ++cc) {
assert((b & extent(ml, rl, cc)).empty());
}
}
}
}
}}}
which b/c of the {{{c}}} and {{cc}} loops is quadratic in the number of
components (MPI ranks).
For many (tens of thousands) components this becomes a dominant cost of
the regrid operation.
The Carpet branch {{{rhaas/quadratic_regrid_time}}} contains a test thorn
ArrayTest and a hacked version of Carpet that can simulate a number of MPI
ranks using just one rank.
One can run:
{{{
mpirun -n 1 exe/cactus_sim arrangements/Carpet/ArrayTest/par/array.par
}}}
and control the number of pretend ranks by setting {{{ArrayTest::size}}}
in array.par.
On my workstation regrid takes ~3.5s for 16k ranks with the consistency
check and ~0.5s without.
This is per grid array and per grid scalar. So for a typical setup with
~400 grid arrays and grid scalars (no matter whether they have storage or
not) this amounts to 400 * 3s = 1200s of time spent in the consistency
check.
This happens only once per simulation (since grid arrays and grid scalars
are only regrid once) but for a test simulation can be quite significant
(at scale).
The branch actually arranges for the consistency check to be skipped if
one defines {{{CARPET_OPTIMISE}}}.
I attach a gnuplot script that takes carpet-timing-statistics.0000.txt and
shows how much time is spent in dh and gh regrid.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/2211>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#2194: Memory increase during regridding
---------------------------------+--------------------
Reporter: wolfgang.kastaun@… | Type: defect
Status: new | Priority: unset
Milestone: | Component: Other
Version: development version | Keywords:
---------------------------------+--------------------
I still have the problem that my BNS runs consume increasing memory during
inspiral, where regridding happens. It limits the runtime of typical BNS
runs to around 1 day, then they are killed by memory exhaustion. After
merger, with fixed grid structure, the problem goes away.
It happens with two different codes, WhiskyThermal+McLachlan and
WhiskyMHD+McLachlan, on all clusters I use (hydra, marconi, supermuc,
fermi, ..), and all releases I used: Payne, Brahe, Tesla, Wheeler.
Regridding controled using CarpetRegrid2.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/2194>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#2184: PAPI does not detect ifort correctly
---------------------------------+-----------------------------------
Reporter: Roland Haas | Type: defect
Status: new | Priority: minor
Milestone: | Component: EinsteinToolkit thorn
Version: development version | Keywords: PAPI
---------------------------------+-----------------------------------
PAPI's makefile attempts to set options for the fortran compiler and
checks for the intel compiler suite by comparing F77 to the string ifort.
This breaks since for Cactus F77 is not ifort but something along the
lines of /long/path/to/ifort . The attached patch replicates the methods
that the Makefile already uses to identify the C compiler to identify the
F77 compiler.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/2184>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#2214: Cactus: handle error in gethostname
---------------------------------+--------------------
Reporter: Roland Haas | Type: defect
Status: new | Priority: minor
Milestone: | Component: Cactus
Version: development version | Keywords:
---------------------------------+--------------------
This works around deficiencies of the gethostname call, namely that it
does not return an error of the supplied buffer is too small and that it
may not include a NUL byte at the end of the buffer if the buffer was too
small and the returned name was truncated.
This should not happen (often) in practise since glibc already has a
similar workaround in place, but may happen eg on a Mac if they indeed use
a non glibc C library.
Pull request is here: https://bitbucket.org/cactuscode/cactus/pull-
requests/55/cactus-handle-error-in-gethostname/diff
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/2214>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit