Present: Frank, Roland, Steve, Ian, Yosef, Peter
Release: * Frank needs (at least) write access to Llama to create the release branches * other release branches and tags were created
Failing ML using tests: * Peter will provide more details * right now it seems as if it really is the constraints that are triggering * initial data and metric data seems fine
Jenkins test systems: * currently have some tests failing and would like to reduce that number to zero * currently no backups in place, need to do backups inside of the VM
Yours, Roland
On 12 Dec 2016, at 17:43, Roland Haas rhaas@aei.mpg.de wrote:
Present: Frank, Roland, Steve, Ian, Yosef, Peter
Release:
- Frank needs (at least) write access to Llama to create the release
branches
This was arranged during the meeting.
- other release branches and tags were created
Failing ML using tests:
- Peter will provide more details
I have created a ticket for this: https://trac.einsteintoolkit.org/ticket/1995. Please add more information to this ticket if you have it! Since the problem goes away when you change optimisation settings from O2 to O1, this looks to me like a compiler bug.
- right now it seems as if it really is the constraints that are
triggering
- initial data and metric data seems fine
Jenkins test systems:
- currently have some tests failing and would like to reduce that
number to zero
The failing tests are
GRHydro.GRHydro_test_shock_weno NaNChecker.nancount
These started failing when we moved the test VM from one host to another, so this is probably CPU-dependent.
- currently no backups in place, need to do backups inside of the VM
We do have backups, but they are not yet automated. Unfortunately, the VM hosting system at NCSA does not provide VM backups, instead expecting users to back up themselves from within the VM. They say there is no funding for providing backups. Hence, we need to invest our own time in implementing a backup system.
-- Ian Hinder http://members.aei.mpg.de/ianhin
Hello all,
The failing tests are
GRHydro.GRHydro_test_shock_weno NaNChecker.nancount
These started failing when we moved the test VM from one host to another, so this is probably CPU-dependent.
NaNCount is odd since the code is new enough that it should just set up nans directly and the cfg file does not set -ffast-math so NaNs should be detected. I will try and see if I can reproduce this on my laptop but it seems this is very specific to the VMs used for Jenkins.
Note that the many failures on comet and gordon seem to be related to file corruption when producing ASCII Carpet output (though CarpetIOASCII seems bug free). The corruption goes away if I force flushing of the buffers after each line (using std::endl instead of "\n"). This happens on both comet and gordon which use different versions of the intel compiler (2015 and 2013) and different gcc compilers (for the C++ library). It almost sounds like a file system issue (they both use similar scratch file systems since they are both at SDSC) and the issue goes away if I run in $HOME rather than $SCRATCH.
Yours, Roland
On Thu, Dec 15, 2016 at 10:05:54AM -0600, Roland Haas wrote:
Note that the many failures on comet and gordon seem to be related to file corruption when producing ASCII Carpet output (though CarpetIOASCII seems bug free). The corruption goes away if I force flushing of the buffers after each line (using std::endl instead of "\n").
The garbled output, if not coming from two processes, look like a race. If it is, an explicit flush after output wouldn't solve the problem, but instead 'just' make it less likely to happen.
This happens on both comet and gordon which use different versions of the intel compiler (2015 and 2013) and different gcc compilers (for the C++ library). It almost sounds like a file system issue (they both use similar scratch file systems since they are both at SDSC) and the issue goes away if I run in $HOME rather than $SCRATCH.
Yes. This smells like the filesystem or OS not invalidating a cache after a (probably cached) write to a file happened. The sysadmins should either know about it, or should be interested in it. If this is indeed the case, it should be possible to generate a simple test case show-casing the problem.
Frank
Hello Frank, all,
The garbled output, if not coming from two processes, look like a race. If it is, an explicit flush after output wouldn't solve the problem, but instead 'just' make it less likely to happen.
It only happens with 2 MPI ranks, yet I have checked that indeed only a single rank writes data (or opens the file). This happens during a single open/close cycle so it also is not related to opening/closing the file.
Yes. This smells like the filesystem or OS not invalidating a cache after a (probably cached) write to a file happened. The sysadmins should either know about it, or should be interested in it. If this is indeed the case, it should be possible to generate a simple test case show-casing the problem.
"Simple" is relative. The simplest reproducer I have is CarpetIOASCII's newsep test which still fires up a lot of Cactus thorns (not really required to just check that "-" is used instead of "::" for file names). A simple test code that writes the same lines of data does not show the problem (when run using MPI and 2 ranks of which only rank 0 actually writes).
Yours, Roland
On Thu, Dec 15, 2016 at 10:27:00AM -0600, Roland Haas wrote:
It only happens with 2 MPI ranks, yet I have checked that indeed only a single rank writes data (or opens the file). This happens during a single open/close cycle so it also is not related to opening/closing the file.
Could it be that, for some reason, two independent simulations are spawned, both writing to the same file?
Frank
Hello Frank,
Could it be that, for some reason, two independent simulations are spawned, both writing to the same file?
That seems not to be the case, I tested for this by having the simulation append the process ID of the writing process to the file name. Only a single file was created.
Yours, Roland
users@lists.einsteintoolkit.org