#1751: [Pull request: CactusUtils/WatchDog] new thorn to automatically terminate
jobs that hang
-----------------------------------+----------------------------------------
Reporter: dradice@… | Owner:
Type: enhancement | Status: new
Priority: unset | Milestone:
Component: EinsteinToolkit thorn | Version: development version
Keywords: |
-----------------------------------+----------------------------------------
WatchDog is thorn that terminates jobs that do not make progress over a
user-defined time frame. Internally, WatchDog updates an internal timer at
CCTK_ANALYSIS and uses the pthread library to spawn watcher thread that
periodically checks if the timer has been updated. If the timer has not
been updated for more than a user-defined time frame, the thread calls
"exit()" to terminate the process (and the job).
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1751>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1765: CarpetRegrid does not check for incorrect outer boundary size
--------------------+-------------------------------------------------------
Reporter: rhaas | Owner: eschnett
Type: defect | Status: new
Priority: unset | Milestone:
Component: Carpet | Version: development version
Keywords: |
--------------------+-------------------------------------------------------
The attached parfile trigger an assert for inconsistent grid structure
inside of CarpetLib. The issue is that boundary_size is set to 1 instead
of 3.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1765>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1761: Problem with building on XC30 CFCA
-----------------------------------+----------------------------------------
Reporter: maxim.barkov@… | Owner:
Type: task | Status: new
Priority: major | Milestone:
Component: Cactus | Version: development version
Keywords: |
-----------------------------------+----------------------------------------
I try to compile ET on CFCA NAOJ cluster Cray XC30.
I used as similar system edison in NERSC.
So far I have problems with build procedure.
CFCA is configured by default for Cray environment.
In configuration file I switch programming enviroment to Intel and correct
paths to actual ones.
Unfortunately simfactory crashed.
{{{
libtool: install: /usr/bin/install -c fftw-wisdom
/home/barkovmm/ET/Cactus/configs/whisky/scratch/external/FFTW3/bin/fftw-
wisdom
Making install in m4
~/ET/Cactus/configs/whisky/scratch
FFTW3: Cleaning up...
FFTW3: Done.
Creating /home/barkovmm/ET/Cactus/configs/whisky/lib/libthorn_FFTW3.a
make: *** [whisky] Error 2
}}}
If I run comand on the cluster
{{{make whisky_11}}}
ET continue configuration and crushed on GSL
{{{
icc: command line warning #10121: overriding '-xCORE-AVX2' with '-xHost'
In file included from results.c(32):
/usr/lib64/gcc/x86_64-suse-linux/4.3/include/varargs.h(4): error: #error
directive: "GCC no longer implements <varargs.h>."
#error "GCC no longer implements <varargs.h>."
^
In file included from results.c(32):
/usr/lib64/gcc/x86_64-suse-linux/4.3/include/varargs.h(5): error: #error
directive: "Revise your code to use <stdarg.h>."
#error "Revise your code to use <stdarg.h>."
^
results.c(95): error #54: too few arguments in invocation of macro
"va_start"
va_start (ap);
^
results.c(95): error: expected an expression
va_start (ap);
^
results.c(155): error #54: too few arguments in invocation of macro
"va_start"
va_start (ap);
^
results.c(155): error: expected an expression
va_start (ap);
^
results.c(232): error #54: too few arguments in invocation of macro
"va_start"
va_start (ap);
^
results.c(232): error: expected an expression
va_start (ap);
^
results.c(309): error #54: too few arguments in invocation of macro
"va_start"
va_start (ap);
^
results.c(309): error: expected an expression
va_start (ap);
^
results.c(364): error #54: too few arguments in invocation of macro
"va_start"
va_start (ap);
^
results.c(364): error: expected an expression
va_start (ap);
^
results.c(408): error #54: too few arguments in invocation of macro
"va_start"
va_start (ap);
^
results.c(408): error: expected an expression
va_start (ap);
^
Internal error: null pointer
compilation aborted for results.c (code 4)
make[6]: *** [results.lo] Error 1
make[5]: *** [all-recursive] Error 1
make[4]: *** [all] Error 2
make[3]: *** [/work/barkovmm/ET/Cactus/configs/whisky_11/scratch/done/GSL]
Error 2
make[2]: *** [make.checked] Error 2
make[1]: ***
[/work/barkovmm/ET/Cactus/configs/whisky_11/lib/libthorn_GSL.a] Error 2
make: *** [whisky_11] Error 2
}}}
Any help or suggestions are very welcome.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1761>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1768: Output of distrib=constant arrays
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Carpet | Version: development version
Keywords: |
-------------------------+--------------------------------------------------
I have a Cactus array with size=1 and distrib=constant. I would like to
see its value on all processes. According to
http://cactuscode.org/pipermail/users/2014-March/003421.html, it has been
observed that CarpetIOASCII (and CarpetIOHDF5) only output the data from
process 0, whereas IOASCII outputs the data from all processes. The
reason I need this is that I want to see the system swap usage on each
process stored in SystemStatistics::process_memory_mb. The data is
legitimately different on each process.
{{{
REAL process_memory_mb TYPE=array DIM=1 SIZE=1 DISTRIB=constant
TAGS='Checkpoint="no"'
}}}
How hard would it be to add support to CarpetIOASCII for outputting grid
array data from all processes for distrib-constant arrays? It already does
this for distrib=default.
To use the current functionality, one way would be to increase the size of
the array and then write code to replicate the data between processes.
However, I don't think there is any way to specify that the array should
have size=NPROCS, so you would have to specify this using a parameter, and
this is a lot of work for something which would more naturally be done in
CarpetIOASCII.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1768>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1767: test system does not remove old test data before running new test
-----------------------+----------------------------------------------------
Reporter: anonymous | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Cactus | Version: development version
Keywords: |
-----------------------+----------------------------------------------------
When running a test, the test system does not delete the old test output
directory first which can hide missing files.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1767>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1766: simfactory creates its log file in repos/log/simfactory.log rather than
Cactus/log/simfactory.log
------------------------+---------------------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version: development version
Keywords: |
------------------------+---------------------------------------------------
right now (I suspect after the switch to git) simfactory creates its log
file in repos/log/simfactory.log (so 2 level upwards from its bin
directory).
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1766>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1211: CarpetIOScalar should write file info for restart files
--------------------+-------------------------------------------------------
Reporter: knarf | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: Carpet | Version:
Keywords: |
--------------------+-------------------------------------------------------
Currently Carpet doesn't write file info for files created by a restarted
simulation (from a checkpoint, writing into a new file). It should instead
write the header if it created the file - which means it need to check
whether it appends or created the file new.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1211>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1282: Enable HDF5 compression by default in Carpet
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: enhancement | Status: new
Priority: major | Milestone:
Component: Carpet | Version:
Keywords: |
-------------------------+--------------------------------------------------
The current default for CarpetIOHDF5::compression_level is 0 (no
compression). I have been using compression in most of my HDF5 files for
years, and have never run into any problem. CPUs are typically much
faster than storage nowadays. I propose that the compression level should
default to 9. This would affect output and checkpoint files, and could
lead to huge space savings. Apart from the checkpoint files being written
and read quicker and taking less disk space, the user should not notice,
as the HDF5 library handles compression transparently.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1282>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1760: Update supported machines wiki page
-------------------------------------+--------------------------------------
Reporter: rhaas | Owner:
Type: task | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit website | Version: development version
Keywords: |
-------------------------------------+--------------------------------------
The wiki page detailed the supported machines
https://docs.einsteintoolkit.org/et-docs/Supported_Machines
should be updated. Ideally this should be done mechanically based on
simfactories machine.ini files.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1760>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit