#2598: thornyflat - production run segmentation fault
Reporter: Maria
Status: new
Milestone:
Version: ET_2021_11
Type: bug
Priority: major
Component:
Comment (by Anuj Kankani):
I wasn’t doing rm -rf configs before, and after doing that everything seems to be working \(I had fixed the typos and \[thornyflat\] before\). I’ve attached the config files I used below. I noticed there were some disabled thorns in the machine file \(i’m assuming something with blas vs openblas?\). I went ahead and kept them disabled since I assume there was a reason for disabling them. Also I had to specify --machine thornyflat during compilation since it does not choose it by default.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2598/thornyflat-produc…
#2598: thornyflat - production run segmentation fault
Reporter: Maria
Status: new
Milestone:
Version: ET_2021_11
Type: bug
Priority: major
Component:
Comment (by Roland Haas):
Oh, I see. Ok that is easier to offer suggestions for. There are two things that come mind.
1. I would try and make sure that simfactory recognizes thornyflat by checking that `./simfactory/bin/sim whoami` returns `thornyflat` and that the mdb entry simfactory uses is the one I expect using `./simfactory/bin/sim print-mdb-entry $(./simfactory/bin/sim whoami | cut -d' ' -f3)`
2. there seems to be a typo in the file name for the submit script. Namely it is called `thornyflay.sub` \(a `y` where it should be `t`\). This could prevent simfactory from finding it.
Are there any warnings when you start compiling a configuration? Sometimes simfactory picks a “default” run script when it cannot find the one specified on the command line or in the machine file :disappointed:
Just to be sure: you tried this compiling a fresh configuration, ideally after a `rm -rf configs` to make sure there are no lingering old files \(a `--reconfig` does not overwrite an existing `RunScript` file in `configs/sim`\)?
The submitscript attached to the ticket does already use `--machine @MACHINE@` which makes sure that simfactory uses the correct machine description when executing the run script.
Note: looking at you file `thornyflat.ini` in the ticket there is a missing `[thornyflat]` that should be at the top of the file and that actually tells simfactory the name of the machine \(the ini file section\).
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2598/thornyflat-produc…
#2598: thornyflat - production run segmentation fault
Reporter: Maria
Status: new
Milestone:
Version: ET_2021_11
Type: bug
Priority: major
Component:
Comment (by Anuj Kankani):
I am actually able to get ETK running on thornyflat \(and spruceknob\) using simfactory by slightly modifying the configuration files attached above . My only issue is that when creating a new simulation, the run and submit script that get used are the generic onces, despite me specifying --machine thornyflat. By manually going in and replacing the run and submit files in the simulation directory, everything works fine. In the thornyflat machine file I specify the thornyflat configuration files, so I’m not sure why it’s still choosing the generic ones. Is there some step I am missing? I have attached the machine file I am using here.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2598/thornyflat-produc…
#2615: Riemann release not compiling with intel 2021.6.0
Reporter: Wolfgang Kastaun
Status: new
Milestone:
Version: ET_2022_05
Type: bug
Priority: minor
Component:
Comment (by Roland Haas):
This was discussed for a bit in today’s ET call. Is there a public cluster where this happens and that could be used for testing? Most the XSEDE systems would be fine as well as SupeMUC at LRZ. A less convenient solution would be a pointer to where one can download Intel’s OneAPI and compiler these days.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2615/riemann-release-n…
#2598: thornyflat - production run segmentation fault
Reporter: Maria
Status: new
Milestone:
Version: ET_2021_11
Type: bug
Priority: major
Component:
Comment (by Roland Haas):
This was discussed in today’s ET call. While difficult to diagnose remotely, the suggestion were that this may be due to mismatched MPI stacks during compile and runtime or due to incorrect `LD_LIBRARY_PATH`. Without access to the system though, this is almost impossible to correctly diagnose.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2598/thornyflat-produc…
#2617: AEILocalInterp can produce GBs of level 1 warnings if interpolation points are outside of the domain
Reporter: Roland Haas
Status: new
Milestone:
Version:
Type: bug
Priority: minor
Component: EinsteinToolkit thorn
AEILocalInterp will produce warning about each point that it is asked to interpolate to that is outside of the domain. This can easily lead to many many essentially identical warnings if eg the center of a spherical surface is set to NaN but a failure of shift tracking.
Instead of producing warnings all the time I would propose the following change in behaviour:
* by default output no more than \(say\) 10 points, then be silent and report “and NN more warnnings” at the end
* add an option to either make the warning level 0 right away \(an error\) or once 10 pints have been reached
the limits would apply to each interpolator call \(LocalInterpUniform call really so technically per grid component\) and thus reset after each call.
Note: AEILocalInterp uses a level 1 warning since it actually does follow documented \(but mostly ignored\) Cactus design that states that a level 1 warning is almost certainly going to lead to incorrect result but that the code can continue \([http://einsteintoolkit.org/referencemanual/ReferenceManual.html#x1-235000A2](http://einsteintoolkit.org/referencemanual/ReferenceManual.html#x1-235000A2)\). So it _is_ indeed using the warning level as documented, just not like everyone else does.
> #define CCTK\_WARN\_ABORT 0 /\* abort the Cactus run \*/
> \#define CCTK\_WARN\_ALERT 1 /\* the results of this run will probably \*/
> /\* be wrong, but this isn’t quite certain, \*/
> /\* so we’re not going to abort the run \*/
> \#define CCTK\_WARN\_COMPLAIN 2 /\* the user should know about this, but \*/
> /\* the results of this run are probably ok \*/
> \#define CCTK\_WARN\_PICKY 3 /\* this is for small problems that can \*/
> /\* probably be ignored, but that careful \*/
> /\* people may want to know about \*/
> \#define CCTK\_WARN\_DEBUG 4 /\* these messages are probably useful \*/
> /\* only for debugging purposes \*/
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2617/aeilocalinterp-ca…
#181: AEILocalInterp should not off-centre the interpolation stencil by default
Reporter: Ian Hinder
Status: open
Milestone:
Version:
Type: bug
Priority: minor
Component: Cactus
Changes (by Roland Haas):
assignee: Roland Haas (was )
responsible: [] (was )
AEILocalInterp by default will off-centre the interpolation stencil if there are insufficient points to perform the interpolation. In a parallel setting, this could happen due to there being insufficient ghost points for the interpolator chosen. The off-centering leads to an interpolation error which is of the correct order but larger than for a centered stencil. More importantly, it leads to different results on different numbers of processes. This violates a basic design principle of Cactus, and the expectation of users, that changing the number of processors should not change the results of a simulation.
I propose that instead of silently off-centering the stencil, AEILocalInterp should abort with an error indicating that there are insufficient ghost-zones. The interpolator options corresponding to this are:
boundary_off_centering_tolerance={0.0 0.0 0.0 0.0 0.0 0.0}
boundary_extrapolation_tolerance={0.0 0.0 0.0 0.0 0.0 0.0}
**Keyword:** AEILocalInterp
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/181/aeilocalinterp-sho…