#2013: piraha breaks ThornDoc building
--------------------+-------------------------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Cactus | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
Doing {{{make ThornDocHTML}}} I get an error message in eg in echo
RELOADAGENT | gpg-connect-agent of
{{{
Undefined subroutine &piraha::parse_peg_file called at
/home/rhaas/postdoc/gr/cactus/ET_trunk/lib/sbin/ScheduleParser.pl line 83.
}}}
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/2013>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1396: Display prompts of executed commands
---------------------------+------------------------------------------------
Reporter: rhaas | Owner: eric9
Type: enhancement | Status: new
Priority: optional | Milestone:
Component: GetComponents | Version: development version
Keywords: |
---------------------------+------------------------------------------------
the attached patch goes to some lengths to capture both stdout and stderr
when executing commands and displays each line of output as it is
generated by the command. This is useful for programs that output prompts
to stdout and wait for user input afterwards.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1396>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#2015: thorn Vectors does not provide a sum() reduction
-------------------------+--------------------------------------------------
Reporter: rhaas | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Cactus | Version: development version
Keywords: Vectors |
-------------------------+--------------------------------------------------
It would be very useful for some applications (eg interpolation) if thorn
Vectors provided a sum() function similar to the sum() member of
vecmathlib
https://bitbucket.org/eschnett/vecmathlib/src/477f9c85cf111e59a9042f230bb53…
=file-view-default#vec_avx_double4.h-537
{{{
real_t sum() const {
// return (*this)[0] + (*this)[1] + (*this)[2] + (*this)[3];
// __m256d x = _mm256_hadd_pd(v, v);
// __m128d xlo = _mm256_extractf128_pd(x, 0);
// __m128d xhi = _mm256_extractf128_pd(x, 1);
realvec_t x = *this;
x = _mm256_hadd_pd(x.v, x.v);
return x[0] + x[2];
}
}}}
Most likely one can just copy and paste the code from vecmathlib.
vecmathlib's license is not LGPL but permissive enough for inclusion and
we can also ask Erik if one can use it and change the license of the
affected code lines to LGPL in Cactus, the license is here:
https://bitbucket.org/eschnett/vecmathlib/src/477f9c85cf111e59a9042f230bb53…
=file-view-default
{{{
Copyright (c) 2012, 2013 Erik Schnetter <eschnetter(a)gmail.com>
Permission is hereby granted, free of charge, to any person obtaining a
copy
of this software and associated documentation files (the "Software"), to
deal
in the Software without restriction, including without limitation the
rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in
all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL
THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING
FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
THE SOFTWARE.
}}}
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/2015>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#2026: param.ccl parser writes (debug?) files v1 and v2 into Cactus root
--------------------+-------------------------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Cactus | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
The file parametre_parser.pl contains
{{{
open($fd,">v1") or die;
print $fd $v1,"\n";
close($fd);
open($fd,">v2") or die;
print $fd $v2,"\n";
close($fd);
}}}
causing it to create files v1 and v2 in the main Cactus root.
Without having looked into this in any more details, this smells like
debug output that should be removed before the release.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/2026>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#68: GetComponents hang due to certificate issues.
--------------------------------+-------------------------------------------
Reporter: diener@… | Owner: eric9
Type: defect | Status: new
Priority: major | Milestone:
Component: GetComponents | Version: ET_2010_06
Keywords: |
--------------------------------+-------------------------------------------
I did a complete new checkout of the EinsteinToolkit on numrel06 using the
version of GetComponents available on the EinsteinToolkit web pages and
had GetComponents
hang with no errors or warnings when checking out
AEIThorns/AEILocalInterp.
Trying the checkout manually I got:
Error validating server certificate for 'https://svn.aei.mpg.de:443':
- The certificate is not issued by a trusted authority. Use the
fingerprint to validate the certificate manually!
Certificate information:
- Hostname: svn.aei.mpg.de
- Valid: from Tue, 23 Feb 2010 16:02:12 GMT until Sun, 22 Feb 2015
16:02:12 GMT
- Issuer: Max-Planck-Gesellschaft, DE
- Fingerprint:
03:b4:e8:6e:d9:09:e9:93:72:e9:ff:fa:df:e4:2c:6d:1d:2a:e4:66
(R)eject, accept (t)emporarily or accept (p)ermanently? p
After accepting the certificate and restarting the checkout proceeded
without problems.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/68>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#181: AEILocalInterp should not off-centre the interpolation stencil by default
----------------------------+-----------------------------------------------
Reporter: hinder | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: Cactus | Version:
Keywords: AEILocalInterp |
----------------------------+-----------------------------------------------
AEILocalInterp by default will off-centre the interpolation stencil if
there are insufficient points to perform the interpolation. In a parallel
setting, this could happen due to there being insufficient ghost points
for the interpolator chosen. The off-centering leads to an interpolation
error which is of the correct order but larger than for a centered
stencil. More importantly, it leads to different results on different
numbers of processes. This violates a basic design principle of Cactus,
and the expectation of users, that changing the number of processors
should not change the results of a simulation.
I propose that instead of silently off-centering the stencil,
AEILocalInterp should abort with an error indicating that there are
insufficient ghost-zones. The interpolator options corresponding to this
are:
boundary_off_centering_tolerance={0.0 0.0 0.0 0.0 0.0 0.0}
boundary_extrapolation_tolerance={0.0 0.0 0.0 0.0 0.0 0.0}
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/181>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#2023: CarpetInterp hangs with openmpi
-----------------------------+----------------------------------------------
Reporter: anonymous | Owner:
Type: defect | Status: new
Priority: unset | Milestone:
Component: Other | Version: development version
Keywords: OpenMPI, Carpet |
-----------------------------+----------------------------------------------
On three occasions with two different simulations (TOV, BNS) the code
hung. In two cases, I could attach a debugger to the running process, and
in both cases the backtrace looked like this:
{{{
#0 0x00007f4ecfd5729a in __GI___pthread_mutex_lock (mutex=0xc40d730) at
../nptl/pthread_mutex_lock.c:79
#1 0x00007f4ec7224807 in ?? () from
/usr/lib/openmpi/lib/openmpi/mca_btl_openib.so
#2 0x00007f4ecdae734a in opal_progress () from /usr/lib/libmpi.so.1
#3 0x00007f4ecda2d3b4 in ompi_request_default_wait_all () from
/usr/lib/libmpi.so.1
#4 0x00007f4ec61cb7b7 in ompi_coll_tuned_sendrecv_actual () from
/usr/lib/openmpi/lib/openmpi/mca_coll_tuned.so
#5 0x00007f4ec61d0df6 in ompi_coll_tuned_alltoallv_intra_pairwise () from
/usr/lib/openmpi/lib/openmpi/mca_coll_tuned.so
#6 0x00007f4ecda3a00f in PMPI_Alltoallv () from /usr/lib/libmpi.so.1
#7 0x00000000015458fe in CarpetInterp::Carpet_DriverInterpolate
(cctkGH_=<optimized out>, N_dims=<optimized out>, local_interp_handle=3,
param_table_handle=6421, coord_system_handle=0, N_interp_points=224,
interp_coords_type_code=130, coords_list=<optimized out>,
N_input_arrays=6, input_array_variable_indices=0x7fff38fbf640,
N_output_arrays=24, output_array_type_codes=0x7fff38fbf660,
output_arrays=0x7fff38fbfb00)
at
/home/wolfgang.kastaun/ET/Payne/Cactus/arrangements/Carpet/CarpetInterp/src/interp.cc:645
#8 0x000000000071ee4f in SymBase_SymmetryInterpolateFaces
(cctkGH_=0xccb1be0, N_dims=<optimized out>, local_interp_handle=3,
param_table_handle=6421, coord_system_handle=0, N_interp_points=224,
interp_coords_type=130, interp_coords=0x7fff38fbe1a0, N_input_arrays=6,
input_array_indices=0x7fff38fbf640, N_output_arrays=24,
output_array_types=0x7fff38fbf660, output_arrays=0x7fff38fbfb00, faces=0)
at
/home/wolfgang.kastaun/ET/Payne/Cactus/arrangements/CactusBase/SymBase/src/Interpolation.c:381
}}}
I'm not sure if this is a problem with Carpet or with my OpenMPI
installation. It only happens rarely, after one day or so.
I'm using ET version Payne, gcc 4.9.2, openmpi 1.6.5 on an infiniband
interconnect.
The code was compiled with OpenMP support and run using 4 threads per
process.
Looking at various variables in Carpet_DriverInterpolate with the
debugger, I noticed that the vector tmp was reported with size -488221088,
but this could also be mis-reported by gdb since the code was compiled
with -O3.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/2023>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#778: Compiling PittNullCode is slow
-----------------------------------+----------------------------------------
Reporter: eschnett | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: EinsteinToolkit thorn | Version:
Keywords: |
-----------------------------------+----------------------------------------
Compiling PittNullCode is very slow on some systems (e.g. with gcc). I
believe this is because files such as NullConstr_R00.F90 contain many
whole-array operations that the compiler has to analyse.
I suggest to rewrite these routines, using e.g. forall or do loops. If we
want to keep the elegant, index-free notation, then I suggest to add an
elemental subroutine for the actual calculations and calling it with whole
arrays.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/778>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1283: Missing data in HDF5 files
--------------------+-------------------------------------------------------
Reporter: hinder | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Cactus | Version:
Keywords: |
--------------------+-------------------------------------------------------
If a simulation runs out of disk space while writing an HDF5 file, the
simulation will terminate. The hdf5 file being written might then be
corrupt, and all data from it may be irretrievable. In that case,
restarting from the last-written checkpoint file may leave a "gap" in the
data corresponding to the period between the start of the failed restart
and the last checkpoint file written.
Steps to reproduce:
• Start a simulation which checkpoints periodically and consists
of several restarts
• Keep all checkpoint files
• Restart 0000 completes successfully and checkpoints at iteration
i1
• Restart 0001 checkpoints once after some evolution at iteration
i2
• Restart 0001 terminates abnormally while writing an HDF5 output
file at iteration i3
• The output file is corrupted and nonrecoverable, so there is no
data from iteration i1 to iteration i3
• Restart 0002 starts at iteration i2 as this is the last
checkpoint available
• The simulation continues until the end, but the data from the
corrupted HDF5 file between iteration i1 and i2 is lost
Possible solutions:
1. Write HDF5 files safely, e.g. by first copying the file to a
new temporary file, performing the write, then atomically moving the
temporary file over the original file. The original file would then
remain in the event of a crash while writing the new file. This could be
very expensive for 3D output files.
2. Start a new set of HDF5 files after each checkpoint. This
seems to be the most efficient and simplest solution, but requires readers
of HDF5 files to be modified to take it into account.
3. Check the consistency of all HDF5 files in the previous
restart(s) on recovery, and recover from the latest checkpoint file for
which all previous HDF5 files are valid. We could use code to check the
HDF5 file, or some other flagging mechanism to indicate that HDF5 writes
were completed successfully; e.g. we could rename the HDF5 file to .tmp
during writes, and rename it back after a successful write. This is
complex and requires Cactus or simfactory to look into previous restarts.
It also only applies to HDF5 files, and requires breaking several
abstraction barriers.
4. Wait for HDF5 journalling support. As far as I know only
metadata journalling is planned, which is probably not enough, and in any
case, they are not actively working on the next version of HDF5 at the
moment due to lack of funding.
5. Checkpoint only on termination of the simulation
In reality, we do not keep all checkpoint files. I usually keep just the
last checkpoint file. I believe that a Cactus simulation will only delete
checkpoint files which it has itself written, which means that there will
generally be one checkpoint file kept per restart; the last one written.
This means that you can always recover from the above situation by
rerunning the restart during which the problem occurred. However, keeping
one checkpoint file per restart is a problem in itself, and we should fix
this as well, which would then mean the potential for losing data in the
case of an interrupted write operation.
Thoughts?
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1283>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1193: CarpetReduce uses lsh for index calculations
--------------------+-------------------------------------------------------
Reporter: knarf | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone: ET_2013_05
Component: Carpet | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
CarpetReduce (e.g. reduce.cc:644) uses lsh to calculate GF indices. This
is wrong in case lsh!=ash. Either use ash or the Cactus macros (which
would probably be even better).
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1193>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit