Hi,
We are down to three failing tests in Jenkins:
SphericalHarmonicRecon.regression_test/2procs SphericalHarmonicReconGen.SpEC-dat-test/2procs SphericalHarmonicReconGen.SpEC-h5-test/2procs
These tests all pass on one process but fail on two processes in Jenkins, which uses the ubuntu.cfg optionlist. They all seem to pass on multiple processes on all other machines (http://einsteintoolkit.org/testsuite_results/index.php), including my laptop with gcc.
The first, SphericalHarmonicRecon.regression_test, fails like this:
WARNING level 0 from host 7ce14e5707a0 process 0 while executing schedule bin NullEvol_Initial, routine NullEvolve::NullEvol_InitialSlice in thorn NullEvolve, file NullEvol_InitialSlice.F90:42: -> Error
The second, SphericalHarmonicReconGen.SpEC-dat-test, fails like this:
NewsB_scri.L02Mm01.asc: substantial differences significant differences on 1 (out of 2) lines maximum absolute difference in column 1 is 963 maximum absolute difference in column 2 is 0.000185770963653907 maximum absolute difference in column 3 is 0.000142466608463344 maximum relative difference in column 1 is 1 maximum relative difference in column 2 is 1 maximum relative difference in column 3 is 1 ...
The third, SphericalHarmonicReconGen.SpEC-h5-test, fails like this:
NewsB_scri.L02Mm01.asc: substantial differences significant differences on 1 (out of 2) lines maximum absolute difference in column 1 is 963 maximum absolute difference in column 2 is 0.000185770963653907 maximum absolute difference in column 3 is 0.000142466608463344 maximum relative difference in column 1 is 1 maximum relative difference in column 2 is 1 maximum relative difference in column 3 is 1 ...
I suspect the second and third failures have the same cause. Do we have any idea why these tests fail? They don't seem to fail on any other machines (http://einsteintoolkit.org/testsuite_results/index.php).
hi Ian,
NewsB_scri.L02Mm01.asc: substantial differences significant differences on 1 (out of 2) lines maximum absolute difference in column 1 is 963 maximum absolute difference in column 2 is 0.000185770963653907 maximum absolute difference in column 3 is 0.000142466608463344 maximum relative difference in column 1 is 1 maximum relative difference in column 2 is 1 maximum relative difference in column 3 is 1 ...
The third, SphericalHarmonicReconGen.SpEC-h5-test, fails like this:
NewsB_scri.L02Mm01.asc: substantial differences significant differences on 1 (out of 2) lines maximum absolute difference in column 1 is 963 maximum absolute difference in column 2 is 0.000185770963653907 maximum absolute difference in column 3 is 0.000142466608463344 maximum relative difference in column 1 is 1 maximum relative difference in column 2 is 1 maximum relative difference in column 3 is 1 ...
I suspect the second and third failures have the same cause. Do we have any idea why these tests fail? They don't seem to fail on any other machines (http://einsteintoolkit.org/testsuite_results/index.php).
i'm not sure whether it's related, but i've had a similar issue with a testsuite of a local thorn i have, which was passing just fine with 1 proc but not with 2 procs (and similar errors). the root of the problem turned out to be that on this particular machine, for whatever reason, when running the parfile with 2 procs the output was "doubled". as if each processor was writing the same thing on the same file. the output itself was the same, but the files that were written were obviously not. so diff was signalling differences in the files where in fact the numbers themselves were (nearly) the same. could you be seeing something like this as well?
cheers, Miguel
On 13 Jul 2017, at 10:48, Miguel Zilhão mzilhao@ffn.ub.es wrote:
hi Ian,
NewsB_scri.L02Mm01.asc: substantial differences significant differences on 1 (out of 2) lines maximum absolute difference in column 1 is 963 maximum absolute difference in column 2 is 0.000185770963653907 maximum absolute difference in column 3 is 0.000142466608463344 maximum relative difference in column 1 is 1 maximum relative difference in column 2 is 1 maximum relative difference in column 3 is 1 ... The third, SphericalHarmonicReconGen.SpEC-h5-test, fails like this: NewsB_scri.L02Mm01.asc: substantial differences significant differences on 1 (out of 2) lines maximum absolute difference in column 1 is 963 maximum absolute difference in column 2 is 0.000185770963653907 maximum absolute difference in column 3 is 0.000142466608463344 maximum relative difference in column 1 is 1 maximum relative difference in column 2 is 1 maximum relative difference in column 3 is 1 ... I suspect the second and third failures have the same cause. Do we have any idea why these tests fail? They don't seem to fail on any other machines (http://einsteintoolkit.org/testsuite_results/index.php).
i'm not sure whether it's related, but i've had a similar issue with a testsuite of a local thorn i have, which was passing just fine with 1 proc but not with 2 procs (and similar errors). the root of the problem turned out to be that on this particular machine, for whatever reason, when running the parfile with 2 procs the output was "doubled". as if each processor was writing the same thing on the same file. the output itself was the same, but the files that were written were obviously not. so diff was signalling differences in the files where in fact the numbers themselves were (nearly) the same. could you be seeing something like this as well?
Hi,
This is in general a problem; Cactus output is not guaranteed to be the same when you change the number of processes. There are ways to work around this in specific cases, which are used for the tests. Specifically, you can use the options
CarpetIOASCII::compact_format = yes CarpetIOASCII::output_ghost_points = no
and look at output only in the z direction. If Carpet splits the grid in the z direction, then I think that for small numbers of processes at least, you are guaranteed that the output will be independent of the number of processes, as the processes will write in component order to the file, and the component ordering is the same as the z-ordering in this case.
In your case, you are probably seeing duplicate points because some points (the ghost points) exist on more than one process, and each process outputs all the points it has. Using output_ghost_points = no might help in that case.
For SphericalHarmonicRecon, the tests use PUGH, not Carpet, and compact_format is not available. I will have to check if this is a diffing issue, but I would be surprised, since it works on other machines on 2 processes.
If this double problem occurs only on a particular machine, then this is likely an MPI problem. If things are misconfigured, you are running two identical serial computations instead of one parallel simulation. The output will be correct, but it will be doubled, and the run will take twice as long.
-erik
On Thu, Jul 13, 2017 at 10:27 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 13 Jul 2017, at 10:48, Miguel Zilhão mzilhao@ffn.ub.es wrote:
hi Ian,
NewsB_scri.L02Mm01.asc: substantial differences significant differences on 1 (out of 2) lines maximum absolute difference in column 1 is 963 maximum absolute difference in column 2 is 0.000185770963653907 maximum absolute difference in column 3 is 0.000142466608463344 maximum relative difference in column 1 is 1 maximum relative difference in column 2 is 1 maximum relative difference in column 3 is 1 ... The third, SphericalHarmonicReconGen.SpEC-h5-test, fails like this: NewsB_scri.L02Mm01.asc: substantial differences significant differences on 1 (out of 2) lines maximum absolute difference in column 1 is 963 maximum absolute difference in column 2 is 0.000185770963653907 maximum absolute difference in column 3 is 0.000142466608463344 maximum relative difference in column 1 is 1 maximum relative difference in column 2 is 1 maximum relative difference in column 3 is 1 ... I suspect the second and third failures have the same cause. Do we have any idea why these tests fail? They don't seem to fail on any other machines (http://einsteintoolkit.org/testsuite_results/index.php).
i'm not sure whether it's related, but i've had a similar issue with a testsuite of a local thorn i have, which was passing just fine with 1 proc but not with 2 procs (and similar errors). the root of the problem turned out to be that on this particular machine, for whatever reason, when running the parfile with 2 procs the output was "doubled". as if each processor was writing the same thing on the same file. the output itself was the same, but the files that were written were obviously not. so diff was signalling differences in the files where in fact the numbers themselves were (nearly) the same. could you be seeing something like this as well?
Hi,
This is in general a problem; Cactus output is not guaranteed to be the same when you change the number of processes. There are ways to work around this in specific cases, which are used for the tests. Specifically, you can use the options
CarpetIOASCII::compact_format = yes CarpetIOASCII::output_ghost_points = no
and look at output only in the z direction. If Carpet splits the grid in the z direction, then I think that for small numbers of processes at least, you are guaranteed that the output will be independent of the number of processes, as the processes will write in component order to the file, and the component ordering is the same as the z-ordering in this case.
In your case, you are probably seeing duplicate points because some points (the ghost points) exist on more than one process, and each process outputs all the points it has. Using output_ghost_points = no might help in that case.
For SphericalHarmonicRecon, the tests use PUGH, not Carpet, and compact_format is not available. I will have to check if this is a diffing issue, but I would be surprised, since it works on other machines on 2 processes.
-- Ian Hinder http://members.aei.mpg.de/ianhin
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Hi Ian, Bela found a bug in the null code and I committed his fix on June 2. That commit included an updated testsuite for SphericalHarmonicRecon, but not for SphericalHarmonicReconGen. I wonder if the SphericalHarmonicReconGen testsuite is out of date.
The failure for SphericalHarmonicRecon is quite weird. The failing code is
if (minval(abs(zeta - dcmplx(1,1))) < 1.0d-10) then Tarr = minloc(abs(stereo_q(:,1)-1.)) loc_q = Tarr(1) Tarr = minloc(abs(stereo_p(1,:)-1.)) loc_p = Tarr(1) if (abs(zeta(loc_q,loc_p) - dcmplx(1.,1.)) .gt. 1d-10) then call CCTK_WARN(0, " Error ") endif
endif I can't figure out why the test is there. However, I also can't figure out how it can fail.
Basically the code uses either two real angular coordinates (q and p) or one complex one (zeta). With zeta = q + i * p. The code is checking that if zeta == 1+i anywhere, that it is equal to 1 + i at the point where q=1 and p=1.
Can you try to output loc_q, loc_p, stereo_p(1,loc_p), stereo_p(loc_q,1), and zeta(loc_q, loq_p)?
On 07/13/2017 04:31 AM, Ian Hinder wrote:
Hi,
We are down to three failing tests in Jenkins:
SphericalHarmonicRecon.regression_test/2procs SphericalHarmonicReconGen.SpEC-dat-test/2procs SphericalHarmonicReconGen.SpEC-h5-test/2procs
These tests all pass on one process but fail on two processes in Jenkins, which uses the ubuntu.cfg optionlist. They all seem to pass on multiple processes on all other machines (http://einsteintoolkit.org/testsuite_results/index.php), including my laptop with gcc.
The first, SphericalHarmonicRecon.regression_test, fails like this:
WARNING level 0 from host 7ce14e5707a0 process 0 while executing schedule bin NullEvol_Initial, routine NullEvolve::NullEvol_InitialSlice in thorn NullEvolve, file NullEvol_InitialSlice.F90:42: -> Error
The second, SphericalHarmonicReconGen.SpEC-dat-test, fails like this:
NewsB_scri.L02Mm01.asc: substantial differences significant differences on 1 (out of 2) lines maximum absolute difference in column 1 is 963 maximum absolute difference in column 2 is 0.000185770963653907 maximum absolute difference in column 3 is 0.000142466608463344 maximum relative difference in column 1 is 1 maximum relative difference in column 2 is 1 maximum relative difference in column 3 is 1 ...
The third, SphericalHarmonicReconGen.SpEC-h5-test, fails like this:
NewsB_scri.L02Mm01.asc: substantial differences significant differences on 1 (out of 2) lines maximum absolute difference in column 1 is 963 maximum absolute difference in column 2 is 0.000185770963653907 maximum absolute difference in column 3 is 0.000142466608463344 maximum relative difference in column 1 is 1 maximum relative difference in column 2 is 1 maximum relative difference in column 3 is 1 ...
I suspect the second and third failures have the same cause. Do we have any idea why these tests fail? They don't seem to fail on any other machines (http://einsteintoolkit.org/testsuite_results/index.php).
-- Ian Hinder http://members.aei.mpg.de/ianhin
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On Thu, Jul 13, 2017 at 11:14:00AM -0400, Yosef Zlochower wrote:
if (minval(abs(zeta - dcmplx(1,1))) < 1.0d-10) then Tarr = minloc(abs(stereo_q(:,1)-1.)) loc_q = Tarr(1) Tarr = minloc(abs(stereo_p(1,:)-1.)) loc_p = Tarr(1) if (abs(zeta(loc_q,loc_p) - dcmplx(1.,1.)) .gt. 1d-10) then call CCTK_WARN(0, " Error ") endif
endifor one complex one (zeta). With zeta = q + i * p. The code is checking that if zeta == 1+i anywhere, that it is equal to 1 + i at the point where q=1 and p=1.
I don't quite understand something about that code. It looks for a location where q is closest to 1, and one where p is closest to i:
minloc(abs(stereo_q(:,1)-1.)) minloc(abs(stereo_p(1,:)-1.))
It then assumes the respective 'other' coordinate is the one it should be looking at:
zeta(loc_q,loc_p)
Is this really always the case (could be, if this comes from some kind of known grid setup, but this is not apparent from the code).
Frank
On 07/13/2017 11:25 AM, Frank Loeffler wrote:
On Thu, Jul 13, 2017 at 11:14:00AM -0400, Yosef Zlochower wrote:
if (minval(abs(zeta - dcmplx(1,1))) < 1.0d-10) then Tarr = minloc(abs(stereo_q(:,1)-1.)) loc_q = Tarr(1) Tarr = minloc(abs(stereo_p(1,:)-1.)) loc_p = Tarr(1) if (abs(zeta(loc_q,loc_p) - dcmplx(1.,1.)) .gt. 1d-10) then call CCTK_WARN(0, " Error ") endif
endifor one complex one (zeta). With zeta = q + i * p. The code is checking that if zeta == 1+i anywhere, that it is equal to 1 + i at the point where q=1 and p=1.
I don't quite understand something about that code. It looks for a location where q is closest to 1, and one where p is closest to i:
minloc(abs(stereo_q(:,1)-1.)) minloc(abs(stereo_p(1,:)-1.))
It then assumes the respective 'other' coordinate is the one it should be looking at:
zeta(loc_q,loc_p)
Is this really always the case (could be, if this comes from some kind of known grid setup, but this is not apparent from the code).
I don't understand the test either, but it should be the case that dble(zeta) = stereo_q and dimag(zeta) = stereo_p
zeta is initialized as zeta = dcmplx(stereo_q,stereo_p) in pittnullcode/NullGrid/src/NullGrid_InitCoord.F90
Frank
On Thu, Jul 13, 2017 at 11:29:12AM -0400, Yosef Zlochower wrote:
On 07/13/2017 11:25 AM, Frank Loeffler wrote:
On Thu, Jul 13, 2017 at 11:14:00AM -0400, Yosef Zlochower wrote:
if (minval(abs(zeta - dcmplx(1,1))) < 1.0d-10) then Tarr = minloc(abs(stereo_q(:,1)-1.)) loc_q = Tarr(1) Tarr = minloc(abs(stereo_p(1,:)-1.)) loc_p = Tarr(1) if (abs(zeta(loc_q,loc_p) - dcmplx(1.,1.)) .gt. 1d-10) then call CCTK_WARN(0, " Error ") endif
endifor one complex one (zeta). With zeta = q + i * p. The code is checking that if zeta == 1+i anywhere, that it is equal to 1 + i at the point where q=1 and p=1.
I don't quite understand something about that code. It looks for a location where q is closest to 1, and one where p is closest to i:
minloc(abs(stereo_q(:,1)-1.)) minloc(abs(stereo_p(1,:)-1.))
It then assumes the respective 'other' coordinate is the one it should be looking at:
zeta(loc_q,loc_p)
Is this really always the case (could be, if this comes from some kind of known grid setup, but this is not apparent from the code).
I don't understand the test either, but it should be the case that dble(zeta) = stereo_q and dimag(zeta) = stereo_p
zeta is initialized as zeta = dcmplx(stereo_q,stereo_p) in pittnullcode/NullGrid/src/NullGrid_InitCoord.F90
If that is all, loc_q would be a location in zeta(:,1) where the real part is close to 1, and loc_p is a location in zeta(1,:) where the imaginary part is close to 1. In general, that wouldn't necessarily mean that at location zeta(loc_q,loc_p) any of the parts would be close to 1, let along both at the same time.
Frank
On 07/13/2017 11:33 AM, Frank Loeffler wrote:
On Thu, Jul 13, 2017 at 11:29:12AM -0400, Yosef Zlochower wrote:
On 07/13/2017 11:25 AM, Frank Loeffler wrote:
On Thu, Jul 13, 2017 at 11:14:00AM -0400, Yosef Zlochower wrote:
if (minval(abs(zeta - dcmplx(1,1))) < 1.0d-10) then Tarr = minloc(abs(stereo_q(:,1)-1.)) loc_q = Tarr(1) Tarr = minloc(abs(stereo_p(1,:)-1.)) loc_p = Tarr(1) if (abs(zeta(loc_q,loc_p) - dcmplx(1.,1.)) .gt. 1d-10) then call CCTK_WARN(0, " Error ") endif
endifor one complex one (zeta). With zeta = q + i * p. The code is checking that if zeta == 1+i anywhere, that it is equal to 1 + i at the point where q=1 and p=1.
I don't quite understand something about that code. It looks for a location where q is closest to 1, and one where p is closest to i:
minloc(abs(stereo_q(:,1)-1.)) minloc(abs(stereo_p(1,:)-1.))
It then assumes the respective 'other' coordinate is the one it should be looking at:
zeta(loc_q,loc_p)
Is this really always the case (could be, if this comes from some kind of known grid setup, but this is not apparent from the code).
I don't understand the test either, but it should be the case that dble(zeta) = stereo_q and dimag(zeta) = stereo_p
zeta is initialized as zeta = dcmplx(stereo_q,stereo_p) in pittnullcode/NullGrid/src/NullGrid_InitCoord.F90
If that is all, loc_q would be a location in zeta(:,1) where the real part is close to 1, and loc_p is a location in zeta(1,:) where the imaginary part is close to 1. In general, that wouldn't necessarily mean that at location zeta(loc_q,loc_p) any of the parts would be close to 1, let along both at the same time.
stereo_q should be constant along the second coordinate and stereo_p along the first.
Frank
On Thu, Jul 13, 2017 at 11:38:28AM -0400, Yosef Zlochower wrote:
stereo_q should be constant along the second coordinate and stereo_p along the first.
Ok, so they do form a grid. Is there some other restriction? Because what the code does essentially is check that there is at least one point closer to 1+i than 1.e-10, but maybe that is just not the case? The code snippet in the mail just looks for the one with the shortest distance.
Frank
On Thu, Jul 13, 2017 at 10:46:33AM -0500, Frank Loeffler wrote:
On Thu, Jul 13, 2017 at 11:38:28AM -0400, Yosef Zlochower wrote:
stereo_q should be constant along the second coordinate and stereo_p along the first.
Ok, so they do form a grid. Is there some other restriction? Because what the code does essentially is check that there is at least one point closer to 1+i than 1.e-10, but maybe that is just not the case? The code snippet in the mail just looks for the one with the shortest distance.
I take this back. I forgot about the following line within which all the other code is:
if (minval(abs(zeta - dcmplx(1,1))) < 1.0d-10) then
This is indeed a little strange. Possibly some round-off or inproper grid setup?
Frank
On 07/13/2017 11:49 AM, Frank Loeffler wrote:
On Thu, Jul 13, 2017 at 10:46:33AM -0500, Frank Loeffler wrote:
On Thu, Jul 13, 2017 at 11:38:28AM -0400, Yosef Zlochower wrote:
stereo_q should be constant along the second coordinate and stereo_p along the first.
Ok, so they do form a grid. Is there some other restriction? Because what the code does essentially is check that there is at least one point closer to 1+i than 1.e-10, but maybe that is just not the case? The code snippet in the mail just looks for the one with the shortest distance.
I take this back. I forgot about the following line within which all the other code is:
if (minval(abs(zeta - dcmplx(1,1))) < 1.0d-10) then
This is indeed a little strange. Possibly some round-off or inproper grid setup?
I worry about the latter. Otherwise I would suggest to just remove the test. It has all the earmarks of some old test to debug some long-forgotten issue that somehow got committed into the official code.
Frank
users@lists.einsteintoolkit.org