On 4 Nov 2010, at 23:55, diener@cct.lsu.edu wrote:
User: diener Date: 2010/11/04 05:55 PM
Modified: /trunk/src/macro/ DXYDG_undefine.h
Log: Undefine the guts instead of the declare. This bug has been there since the beginning and showed up when using ADM with the leapfrog scheme using a predictor-corrector step at the first iteration. The source code to calculate the second derivative of the matric with respect to x and y was not included in the pre-processed source code for the corrector step, resulting in the value calculated for the last point in the predictor step was used for all grid points. This lead to wrong results that depended on the number of processors used, since different values where used on different processors.
Good catch! Does this mean that the testsuite results were wrong all along, and now have to be regenerated? Before this fix, the test_ADM_2 test passed on 1 but not 2 processes, and after the fix it fails on both.
What other codes would have been affected? For example, BSSN_MoL?
Again I repeat my call for correctness tests in addition to regression tests...
Hi,
On Fri, Nov 05, 2010 at 08:41:45AM +0100, Ian Hinder wrote:
Good catch! Does this mean that the testsuite results were wrong all along, and now have to be regenerated?
Yes. They have been committed by now. (There was some permission issue to sort out, which is why it took a bit.)
What other codes would have been affected? For example, BSSN_MoL?
Maybe. Potentially everything which included that file twice in the same file.
Again I repeat my call for correctness tests in addition to regression tests...
I agree. This morning we found another testsuite which has completely wrong result data and Roland is about to fix this.
Frank
On 5 Nov 2010, at 17:13, Frank Loeffler wrote:
Hi,
On Fri, Nov 05, 2010 at 08:41:45AM +0100, Ian Hinder wrote:
Good catch! Does this mean that the testsuite results were wrong all along, and now have to be regenerated?
Yes. They have been committed by now. (There was some permission issue to sort out, which is why it took a bit.)
What other codes would have been affected? For example, BSSN_MoL?
Maybe. Potentially everything which included that file twice in the same file.
Again I repeat my call for correctness tests in addition to regression tests...
I agree. This morning we found another testsuite which has completely wrong result data and Roland is about to fix this.
Which one was that?
I have also noticed that the WeylScal4 test is newly failing as of last night. I modified the parameter files yesterday to remove the Carpet::poison_value parameter as the new Mercurial version of Carpet doesn't understand that parameter. The test passed on my laptop so I committed the new parameter file, but it seems to fail with NaNs on Damiana. I'm looking into it now. It could also theoretically have been caused by some other change that happened yesterday.
On Fri, Nov 05, 2010 at 05:32:49PM +0100, Ian Hinder wrote:
I have also noticed that the WeylScal4 test is newly failing as of last night.
This one.
I modified the parameter files yesterday to remove the Carpet::poison_value parameter as the new Mercurial version of Carpet doesn't understand that parameter. The test passed on my laptop so I committed the new parameter file, but it seems to fail with NaNs on Damiana. I'm looking into it now.
Roland is as well. The problem doesn't seem to come from Weylscal4 itself, because already the evolved metric develops NaNs. Looking at the used refinement levels I would guess that they are simply too small, without enough space between them.
Frank
On 5 Nov 2010, at 17:40, Frank Loeffler wrote:
On Fri, Nov 05, 2010 at 05:32:49PM +0100, Ian Hinder wrote:
I have also noticed that the WeylScal4 test is newly failing as of last night.
This one.
I modified the parameter files yesterday to remove the Carpet::poison_value parameter as the new Mercurial version of Carpet doesn't understand that parameter. The test passed on my laptop so I committed the new parameter file, but it seems to fail with NaNs on Damiana. I'm looking into it now.
Roland is as well. The problem doesn't seem to come from Weylscal4 itself, because already the evolved metric develops NaNs. Looking at the used refinement levels I would guess that they are simply too small, without enough space between them.
That sounds likely. It's probably a good idea to use NaN as the poison value (which is the default) in all test suites. For production runs, it's useful to have an identifiable non-NaN value because (a) it propagates more slowly, meaning that it can be seen when visualising data output every few iterations, and (b) because you can tell that it comes from a bug rather than a numerical problem. But in a test suite, we want any failure to be as obvious as possible, and a NaN is a good way to get a problem noticed. For example, any norm output would be sensitive to it.
On Fri, Nov 05, 2010 at 05:48:51PM +0100, Ian Hinder wrote:
But in a test suite, we want any failure to be as obvious as possible, and a NaN is a good way to get a problem noticed.
I agree.
Except of course in the very special testsuite which would test the poison values itself. :)
Frank
Hi,
On Fri, 5 Nov 2010, Frank Loeffler wrote:
Hi,
On Fri, Nov 05, 2010 at 08:41:45AM +0100, Ian Hinder wrote:
Good catch! Does this mean that the testsuite results were wrong all along, and now have to be regenerated?
Yes. They have been committed by now. (There was some permission issue to sort out, which is why it took a bit.)
What other codes would have been affected? For example, BSSN_MoL?
Maybe. Potentially everything which included that file twice in the same file.
BSSN_MoL would not have been affected since it uses it's own macros and also doesn't uses derivatives of the adm metric variables. As Frank says the only thorns that would have been affected by this would have had to include this particular macro twice in the same file. This doesn't seem to be the case for ADMConstraints for example.
Again I repeat my call for correctness tests in addition to regression tests...
I agree. This morning we found another testsuite which has completely wrong result data and Roland is about to fix this.
This bug has been around since the first commit of ADMMacros in 1999, so I don't quite understand how it failed to get noticed. I guess nobody ever used this particular evolution scheme in production. We did notice it for the first release of the Einstein Toolkit but at the time we decided that since ADM wasn't used by anybody it wasn't worth spending time investigating it.
Cheers,
Peter
users@lists.einsteintoolkit.org