Hi,
Bruno Mundim will chair next weeks phone call because I will be traveling at that time. We do have a release to finish and I wouldn't like to wait for another week for that to happen - besides I will be traveling next week as well (but might be able to call in).
The call will, as usual, happen at 10am central US time (now winter time). Details about how to connect are available at http://einsteintoolkit.org/info/support/ .
The next release of the Einstein Toolkit will be the Nr. 1 item on the agenda of the next call tomorrow. We should all try to concentrate for a moment on the release, get it out of the door, to then be able to go on with our usual great science without this on the back of our minds.
However, first we need to look at a few outstanding issues:
1) Testsuites A lot has been done over the course of last week, and I would like to thank all involved people. A few issues remain however: 1a) Most AHFinderDirect testsuites fail on 2 processes. This is not new. This has been the same for last release. However, in contrast to some other thorns, AHFinderDirect is used by almost every group, in production simulations. We should really figure out what goes wrong here. Even more severe: preliminary tests suggest that this could be connected to the interpolator, which would make matters only worse if it turns out to be true. 1b) Some GRHydro testsuites (especially involving ppm) fail on some machines. This is possibly simply connected to numerical errors. The deviations are small, while the default relative tolerance for con2prim is higher than that. Before cranking up the tolerances for the testsuites we should be sure about that though. 1c) McLachlan testsuites fail on the (potentially) three largest machines: ranger, kraken and bluedrop, but pass everywhere else. 1d) WeylScal4/teukolsky fails. The origin of this is known and simply a bug in the testsuite setup itself. Either we find time to correct this (is Roland still on it?), or we drop the testsuite. I ordered the list above by my personal feeling of importance. We could release with all of them unfixed (except 1d), since nothing of this is new, but I would really prefer otherwise.
Please concentrate on the testsuites until Wednesday, so that we can prepare the release for the end of the week.
2) Bruno did some work on the web page and documentation and will probably report about the progress there.
3) The (short) release text is not yet on the wiki, but some of the details in the long text have been added. Please expand this with details of changes since the last release.
Frank Loeffler
Hello all,
1c) McLachlan testsuites fail on the (potentially) three largest machines: ranger, kraken and bluedrop, but pass everywhere else. 1d) WeylScal4/teukolsky fails. The origin of this is known and simply a bug in the testsuite setup itself. Either we find time to correct this (is Roland still on it?), or we drop the testsuite.
I am, it goes away if I increase the number of ghost zones (though I don't understand why). I'll send a parameter file to Ian to test. Without increasing the ghost zones the error persists even when evolving Minkowski spacetime and with much larger grids.
Yours, Roland
Hi,
- Testsuites
A lot has been done over the course of last week, and I would like to thank all involved people. A few issues remain however: 1a) Most AHFinderDirect testsuites fail on 2 processes. This is not new. This has been the same for last release. However, in contrast to some other thorns, AHFinderDirect is used by almost every group, in production simulations. We should really figure out what goes wrong here. Even more severe: preliminary tests suggest that this could be connected to the interpolator, which would make matters only worse if it turns out to be true.
I took a bit closer look at this trying to figure out if the problem is with the interpolator or not. Defining in gr/expansion.cc:
#define GEOMETRY_INTERP_DEBUG2
and changing the output formatting from using %g to %15.11g to get more digits in the debug output it was clear that there was differences in the interpolation results (of the order 1e-5) when running the test parameter file Kerr-definition-expansion.par on 1 or 2 processors, even though the interpolation points were identical.
That parameter file uses the interpolation parameters:
AHFinderDirect::geometry_interpolator_name = "Lagrange polynomial interpolation" AHFinderDirect::geometry_interpolator_pars = "order=4"
I see the same when using:
AHFinderDirect::geometry_interpolator_name = "Lagrange polynomial interpolation" AHFinderDirect::geometry_interpolator_pars = "order=2"
while the following combinations all give identical bh masses on 1 or 2 processors:
AHFinderDirect::geometry_interpolator_name = "Lagrange polynomial interpolation" AHFinderDirect::geometry_interpolator_pars = "order=3"
AHFinderDirect::geometry_interpolator_name = "Hermite polynomial interpolation" AHFinderDirect::geometry_interpolator_pars = "order=2"
AHFinderDirect::geometry_interpolator_name = "Hermite polynomial interpolation" AHFinderDirect::geometry_interpolator_pars = "order=3"
In those cases the largest differences between 1 and 2 processor runs seem to be of order 1e-13.
So it seems there is a bug in AEILocalInterp in the 4th and 2nd order Lagrange interpolation schemes, while the 3rd order Lagrange interpolation and 2nd and 3rd order Hermite interpolation schemes are okay.
Does anybody feel up to looking into this?
Cheers,
Peter
I tested further, using Carpet instead of PUGH. I believe that the problem is in the single processor case of PUGHInterp, i.e. that the 2 process results are actually correct. Carpet gives identical answers for 1 and 2 processes.
-erik
On Mon, Nov 8, 2010 at 8:03 AM, Peter Diener diener@cct.lsu.edu wrote:
Hi,
- Testsuites
A lot has been done over the course of last week, and I would like to thank all involved people. A few issues remain however: 1a) Most AHFinderDirect testsuites fail on 2 processes. This is not new. This has been the same for last release. However, in contrast to some other thorns, AHFinderDirect is used by almost every group, in production simulations. We should really figure out what goes wrong here. Even more severe: preliminary tests suggest that this could be connected to the interpolator, which would make matters only worse if it turns out to be true.
I took a bit closer look at this trying to figure out if the problem is with the interpolator or not. Defining in gr/expansion.cc:
#define GEOMETRY_INTERP_DEBUG2
and changing the output formatting from using %g to %15.11g to get more digits in the debug output it was clear that there was differences in the interpolation results (of the order 1e-5) when running the test parameter file Kerr-definition-expansion.par on 1 or 2 processors, even though the interpolation points were identical.
That parameter file uses the interpolation parameters:
AHFinderDirect::geometry_interpolator_name = "Lagrange polynomial interpolation" AHFinderDirect::geometry_interpolator_pars = "order=4"
I see the same when using:
AHFinderDirect::geometry_interpolator_name = "Lagrange polynomial interpolation" AHFinderDirect::geometry_interpolator_pars = "order=2"
while the following combinations all give identical bh masses on 1 or 2 processors:
AHFinderDirect::geometry_interpolator_name = "Lagrange polynomial interpolation" AHFinderDirect::geometry_interpolator_pars = "order=3"
AHFinderDirect::geometry_interpolator_name = "Hermite polynomial interpolation" AHFinderDirect::geometry_interpolator_pars = "order=2"
AHFinderDirect::geometry_interpolator_name = "Hermite polynomial interpolation" AHFinderDirect::geometry_interpolator_pars = "order=3"
In those cases the largest differences between 1 and 2 processor runs seem to be of order 1e-13.
So it seems there is a bug in AEILocalInterp in the 4th and 2nd order Lagrange interpolation schemes, while the 3rd order Lagrange interpolation and 2nd and 3rd order Hermite interpolation schemes are okay.
Does anybody feel up to looking into this?
Cheers,
Peter
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On Nov 8, 2010, at 5:58 PM, Erik Schnetter wrote:
I tested further, using Carpet instead of PUGH. I believe that the problem is in the single processor case of PUGHInterp, i.e. that the 2 process results are actually correct. Carpet gives identical answers for 1 and 2 processes.
Interesting, we concurrently did the same test, but I still did get differences. Can you send me your parameter file?
Thanks, Eloisa
On Mon, Nov 8, 2010 at 12:07 PM, Eloisa Bentivegna bentivegna@cct.lsu.eduwrote:
On Nov 8, 2010, at 5:58 PM, Erik Schnetter wrote:
I tested further, using Carpet instead of PUGH. I believe that the
problem is in the single processor case of PUGHInterp, i.e. that the 2 process results are actually correct. Carpet gives identical answers for 1 and 2 processes.
Interesting, we concurrently did the same test, but I still did get differences. Can you send me your parameter file?
I used the Mercurial version of Carpet, and I disabled bitant mode. I attach my parameter file.
-erik
Hi,
A bit more information to add to the confusion. Using the git version of Carpet, I was not able to get the same results on 1 and 2 processors. Not even going to a full grid as Erik did. However, with the added interpolation debug output built into AHFinderDirect I do notice a somwehat different behavior. With PUGH I see differences (within 11-12 decimal places) in the interpolation results before the first call to the linear solver. With Carpet I don't. There the first differences appear in the horizon coordinates after the first linear solve. So with Carpet it appears that the linear solver gets the same data on 1 and 2 procs but finds a different solution.
I get the same results using PUGH and Carpet on 1 processor. The results on 2 processors using PUGH and Carpet differs from the 1 processor results and with each other.
Also the results differ depending on whether bitant or full mode is used even on 1 processor.
Cheers,
Peter
On Mon, 8 Nov 2010, Erik Schnetter wrote:
On Mon, Nov 8, 2010 at 12:07 PM, Eloisa Bentivegna bentivegna@cct.lsu.edu wrote: On Nov 8, 2010, at 5:58 PM, Erik Schnetter wrote:
> I tested further, using Carpet instead of PUGH. I believe that the problem is in the single processor case of PUGHInterp, i.e. that the 2 process results are actually correct. Carpet gives identical answers for 1 and 2 processes.Interesting, we concurrently did the same test, but I still did get differences. Can you send me your parameter file?
I used the Mercurial version of Carpet, and I disabled bitant mode. I attach my parameter file.
-erik
-- Erik Schnetter schnetter@cct.lsu.edu http://www.cct.lsu.edu/~eschnett/
On Nov 9, 2010, at 8:30 AM, Peter Diener wrote:
Hi,
A bit more information to add to the confusion. Using the git version of Carpet, I was not able to get the same results on 1 and 2 processors. Not even going to a full grid as Erik did. However, with the added interpolation debug output built into AHFinderDirect I do notice a somwehat different behavior. With PUGH I see differences (within 11-12 decimal places) in the interpolation results before the first call to the linear solver. With Carpet I don't. There the first differences appear in the horizon coordinates after the first linear solve. So with Carpet it appears that the linear solver gets the same data on 1 and 2 procs but finds a different solution.
Peter,
by debug output, do you mean that you're also checking mean curvature and expansion, iteration by iteration? The linear solver also needs these quantities, along with the shape h, to compute the next surface guess. Since both mean curvature and expansion require interpolation, this could explain the differences between different processor counts, even if the initial guesses for h are identical.
I get the same results using PUGH and Carpet on 1 processor. The results on 2 processors using PUGH and Carpet differs from the 1 processor results and with each other.
Also the results differ depending on whether bitant or full mode is used even on 1 processor.
I think here it is important to state whether the differences are compatible with roundoff or not. Identical results are guaranteed only for identical operations, and turning symmetries on and off does not generally satisfy this criterion.
I've checked my tests and I do find that, with the git version of Carpet, the differences reduce from about 1e-06 to 1e-14 by enabling full mode. The latter is arguably acceptable, the former obviously not.
Eloisa
On Tue, Nov 9, 2010 at 2:30 AM, Peter Diener diener@cct.lsu.edu wrote:
Also the results differ depending on whether bitant or full mode is used
even on 1 processor.
AHFinderDirect uses a different grid when using CartGrid3D's bitant mode. This can explain the different solution found by the elliptic solver between bitant and full mode.
-erik
users@lists.einsteintoolkit.org