Hi all,
I have written some test cases for the McLachlan evolution code. This thorn is currently used in other test suites, but it doesn't have one of its own. The minimal test I have written is a 3D diagonal gauge wave in unigrid with periodic boundary conditions run for one crossing time, so mostly it tests that the RHSs are correct. There are two resolutions (15^3 and 30^3) as well as the exact solution as computed by Exact, so from the test suite output you can compute convergence to test that everything is correct (the automatic regression testsuite mechanism does not do this though). I attach the parameter files for comments. I have tested 4th order convergence to the exact solution.
I have enabled 3D HDF5 output in the test suite in anticipation that we might want to have HDF5 comparisons added at some point, and so that when investigating failures there is full information available. The output is from the initial data and the last iteration only, with full compression.
The test output sizes are:
n = 15 2.9 M n = 30 7.9 M Exact 100 kB
Total: 11 M
For the exact solution I only output 3D of the admbase gxx and kxx variables. If this is considered too large, we can omit the 3D HDF5 output for now, as it is not used in the current Cactus testsuite mechanism.
I have put the test suites in the existing ML_BSSN_Helper thorn, as this seemed the easiest solution, since ML_BSSN is automatically generated.
The test does not test the shift terms. For that I would use the shifted gauge wave solution, but as far as I know the harmonic shift gauge condition is not implemented in McLachlan.
The tests are run on 2 processes as CarpetIOASCII output is used, so the output depends on the number of processes. If output on 1 process is better, let me know. In principle, the number of OpenMP threads should not matter in this case. However, the tests seem to fail if I use OMP_NUM_THREADS=2 (or leave it unset on my dual-core laptop, which makes it use 2). The failure is:
Issuing mpirun -np 2 /Users/ian/Cactus/llama/exe/cactus_mcl /Users/ ian/Cactus/llama/arrangements/McLachlan/ML_BSSN_Helper/test/ gw3d_Exact_ord4_15.par
alp.norm1.asc: differences below tolerance on 1 lines alp.norm2.asc: differences below tolerance on 1 lines alp.sum.asc: substantial differences significant differences on 1 (out of 2) lines maximum absolute difference in column 3 is 1.00044417195022e-11 maximum relative difference in column 3 is 2.96613722781153e-15
Failure: 12 files compared, 3 differ, 1 differ significantly
Files which differ strongly: [1] alp.sum.asc
I don't know why this fails.
On Jun 14, 2010, at 11:07 , Ian Hinder wrote:
Hi all,
I have written some test cases for the McLachlan evolution code.
Yay! Thanks!
This thorn is currently used in other test suites, but it doesn't have one of its own. The minimal test I have written is a 3D diagonal gauge wave in unigrid with periodic boundary conditions run for one crossing time, so mostly it tests that the RHSs are correct. There are two resolutions (15^3 and 30^3) as well as the exact solution as computed by Exact, so from the test suite output you can compute convergence to test that everything is correct (the automatic regression testsuite mechanism does not do this though). I attach the parameter files for comments. I have tested 4th order convergence to the exact solution.
I have enabled 3D HDF5 output in the test suite in anticipation that we might want to have HDF5 comparisons added at some point, and so that when investigating failures there is full information available. The output is from the initial data and the last iteration only, with full compression.
The test output sizes are:
n = 15 2.9 M n = 30 7.9 M Exact 100 kB
Total: 11 M
For the exact solution I only output 3D of the admbase gxx and kxx variables. If this is considered too large, we can omit the 3D HDF5 output for now, as it is not used in the current Cactus testsuite mechanism.
I would omit them. I did include a lot of test result output in IsolatedHorizon, and I received many complaints from people about the time it takes to download the thorn, and the disk space it uses. I don't really agree with these sentiments, but I now think that extensive output should be stored somewhere else -- maybe in a special repository that one checks out only when test cases fail.
I have put the test suites in the existing ML_BSSN_Helper thorn, as this seemed the easiest solution, since ML_BSSN is automatically generated.
Actually, the ML_BSSN_Helper thorns are also generated automatically since they are all very similar. Their source is in the "prototype" subdirectory. This process would overwrite your test cases. You could introduce a new thorn, e.g. ML_BSSN_Test.
The test does not test the shift terms. For that I would use the shifted gauge wave solution, but as far as I know the harmonic shift gauge condition is not implemented in McLachlan.
The tests are run on 2 processes as CarpetIOASCII output is used, so the output depends on the number of processes. If output on 1 process is better, let me know.
I prefer output on 2 processors.
In principle, the number of OpenMP threads should not matter in this case. However, the tests seem to fail if I use OMP_NUM_THREADS=2 (or leave it unset on my dual-core laptop, which makes it use 2). The failure is:
Issuing mpirun -np 2 /Users/ian/Cactus/llama/exe/cactus_mcl /Users/ ian/Cactus/llama/arrangements/McLachlan/ML_BSSN_Helper/test/ gw3d_Exact_ord4_15.par
alp.norm1.asc: differences below tolerance on 1 lines alp.norm2.asc: differences below tolerance on 1 lines alp.sum.asc: substantial differences significant differences on 1 (out of 2) lines maximum absolute difference in column 3 is 1.00044417195022e-11 maximum relative difference in column 3 is 2.96613722781153e-15
Failure: 12 files compared, 3 differ, 1 differ significantly
Files which differ strongly: [1] alp.sum.asc
I don't know why this fails.
With N grid points, the absolute error in the sum reduction is N times larger than the absolute error in the average reduction. I usually disable this reduction operation (and some others) in test cases.
-erik
On 16 Jun 2010, at 04:34, Erik Schnetter wrote:
On Jun 14, 2010, at 11:07 , Ian Hinder wrote:
Hi all,
I have written some test cases for the McLachlan evolution code.
Yay! Thanks!
For the exact solution I only output 3D of the admbase gxx and kxx variables. If this is considered too large, we can omit the 3D HDF5 output for now, as it is not used in the current Cactus testsuite mechanism.
I would omit them. I did include a lot of test result output in IsolatedHorizon, and I received many complaints from people about the time it takes to download the thorn, and the disk space it uses. I don't really agree with these sentiments, but I now think that extensive output should be stored somewhere else -- maybe in a special repository that one checks out only when test cases fail.
I agree - I also don't like to have to download a large amount of testsuite data, or have it taking up home directory quota space. If we have 200 thorns, and each of them are using 10 MB for testsuites, that is 2 GB per source tree.
I have modified the tests to output only 1D ASCII data in the x, y and z directions. This should be enough to catch regressions, since all the points on the grid will couple to each other in a single crossing time. The ML_BSSN_Test thorn is now 1.4 MB uncompressed (125 KB compressed)
I have put the test suites in the existing ML_BSSN_Helper thorn, as this seemed the easiest solution, since ML_BSSN is automatically generated.
Actually, the ML_BSSN_Helper thorns are also generated automatically since they are all very similar. Their source is in the "prototype" subdirectory. This process would overwrite your test cases. You could introduce a new thorn, e.g. ML_BSSN_Test.
I have done this now.
The test does not test the shift terms. For that I would use the shifted gauge wave solution, but as far as I know the harmonic shift gauge condition is not implemented in McLachlan.
The tests are run on 2 processes as CarpetIOASCII output is used, so the output depends on the number of processes. If output on 1 process is better, let me know.
I prefer output on 2 processors.
Good.
In principle, the number of OpenMP threads should not matter in this case. However, the tests seem to fail if I use OMP_NUM_THREADS=2 (or leave it unset on my dual-core laptop, which makes it use 2). The failure is:
Issuing mpirun -np 2 /Users/ian/Cactus/llama/exe/cactus_mcl /Users/ ian/Cactus/llama/arrangements/McLachlan/ML_BSSN_Helper/test/ gw3d_Exact_ord4_15.par
alp.norm1.asc: differences below tolerance on 1 lines alp.norm2.asc: differences below tolerance on 1 lines alp.sum.asc: substantial differences significant differences on 1 (out of 2) lines maximum absolute difference in column 3 is 1.00044417195022e-11 maximum relative difference in column 3 is 2.96613722781153e-15
Failure: 12 files compared, 3 differ, 1 differ significantly
Files which differ strongly: [1] alp.sum.asc
I don't know why this fails.
With N grid points, the absolute error in the sum reduction is N times larger than the absolute error in the average reduction. I usually disable this reduction operation (and some others) in test cases.
I have removed output for the reduction, but I now get differences above tolerance in the Hamiltonian constraint:
H.x.asc: substantial differences significant differences on 6 (out of 46) lines maximum absolute difference in column 13 is 3.46657468355827e-12 maximum relative difference in column 13 is 9.69181535413736e-11 (insignificant differences on 9 lines) H.y.asc: substantial differences significant differences on 7 (out of 46) lines maximum absolute difference in column 13 is 2.2119042708546e-12 maximum relative difference in column 13 is 4.60112943759897e-11 (insignificant differences on 8 lines) H.z.asc: substantial differences significant differences on 14 (out of 60) lines maximum absolute difference in column 13 is 3.00423574906006e-12 maximum relative difference in column 13 is 6.2493109449911e-11 (insignificant differences on 7 lines)
The evolved variables show differences below tolerance.
I attach the patch which contains the testsuites. I am inclined to commit it, and warn people to use OMP_NUM_THREADS=1 when running it. At least this gives us a direct working regression test for McLachlan. But we should understand why the Hamiltonian constraint is giving these differences.
If we want to include this test in the release, we will need to modify the einsteintoolkit.th thornlist to include the new thorn. The test runs only on 2 processes, as specified in the test.ccl file, due to the CarpetIOASCII output being dependent on the process decomposition. I have run the test successfully on my laptop and on damiana, but not on anything more exotic.
On 16 Jun 2010, at 16:11, Ian Hinder wrote:
I have removed output for the reduction, but I now get differences above tolerance in the Hamiltonian constraint:
H.x.asc: substantial differences significant differences on 6 (out of 46) lines maximum absolute difference in column 13 is 3.46657468355827e-12 maximum relative difference in column 13 is 9.69181535413736e-11 (insignificant differences on 9 lines) H.y.asc: substantial differences significant differences on 7 (out of 46) lines maximum absolute difference in column 13 is 2.2119042708546e-12 maximum relative difference in column 13 is 4.60112943759897e-11 (insignificant differences on 8 lines) H.z.asc: substantial differences significant differences on 14 (out of 60) lines maximum absolute difference in column 13 is 3.00423574906006e-12 maximum relative difference in column 13 is 6.2493109449911e-11 (insignificant differences on 7 lines)
The evolved variables show differences below tolerance.
I attach the patch which contains the testsuites. I am inclined to commit it, and warn people to use OMP_NUM_THREADS=1 when running it. At least this gives us a direct working regression test for McLachlan. But we should understand why the Hamiltonian constraint is giving these differences.
If we want to include this test in the release, we will need to modify the einsteintoolkit.th thornlist to include the new thorn. The test runs only on 2 processes, as specified in the test.ccl file, due to the CarpetIOASCII output being dependent on the process decomposition. I have run the test successfully on my laptop and on damiana, but not on anything more exotic.
Erik agreed that we should commit the testsuite, so I have committed it. I attach a patch to add the test thorn to the einsteintoolkit.th thornlist. OK to commit?
On Wed, Jun 16, 2010 at 04:11:20PM +0200, Ian Hinder wrote:
I have removed output for the reduction, but I now get differences above tolerance in the Hamiltonian constraint:
I see the same, for OMP_NUM_THREADS=1 and two mpi processes. That level or errors in the constraint is probably to be expected and we should increase the tolerance value for those variables. However, we currently cannot do that on a per variable-basis and cannot implement and test this to be ready for the release. Thus, I suggest to split the testsuite for the moment into two almost identical suites, one generating only the contraint output, and one generating all other output. Those two can then have different tolerance levels. Once we implemented something in Cactus which can set the tolerance on a per-file basis, we could combine those again.
Frank
On 16 Jun 2010, at 20:26, Frank Loeffler wrote:
On Wed, Jun 16, 2010 at 04:11:20PM +0200, Ian Hinder wrote:
I have removed output for the reduction, but I now get differences above tolerance in the Hamiltonian constraint:
I see the same, for OMP_NUM_THREADS=1 and two mpi processes. That level or errors in the constraint is probably to be expected and we should increase the tolerance value for those variables.
I was thinking, very naively, that since the Hamiltonian constraint contains 2nd derivatives of the metric, when dividing by dx^2, we would expect the roundoff errors to be amplified by about two orders of magnitude. So if we are just barely meeting the default 1e-11 Cactus tolerance in the evolved variables, then the constraints might be worse. But I don't really understand this.
However, I have just tried using the Intel compiler options "-fp-model precise -fp-model source" as recommended by Intel if you want reproducibility of floating point computations,
http://software.intel.com/en-us/articles/consistency-of-floating-point-resul...
and if you do that, the test suites pass with "files identical" with 1, 2 and 8 threads on my laptop. So this confirms that the problem comes from a behaviour of the compiler, and not from a problem in the code. Of course we would like to test the compiler settings we actually use in production, and the above settings probably cause a significant loss of optimisation.
However, we currently cannot do that on a per variable-basis and cannot implement and test this to be ready for the release. Thus, I suggest to split the testsuite for the moment into two almost identical suites, one generating only the contraint output, and one generating all other output. Those two can then have different tolerance levels. Once we implemented something in Cactus which can set the tolerance on a per-file basis, we could combine those again.
Is it right to set the tolerance on what we observe, rather than on what we conclude we should expect? It feels a bit like cheating to me.
Hi Ian,
On Thu, 17 Jun 2010, Ian Hinder wrote:
On 16 Jun 2010, at 20:26, Frank Loeffler wrote:
On Wed, Jun 16, 2010 at 04:11:20PM +0200, Ian Hinder wrote:
I have removed output for the reduction, but I now get differences above tolerance in the Hamiltonian constraint:
I see the same, for OMP_NUM_THREADS=1 and two mpi processes. That level or errors in the constraint is probably to be expected and we should increase the tolerance value for those variables.
I was thinking, very naively, that since the Hamiltonian constraint contains 2nd derivatives of the metric, when dividing by dx^2, we would expect the roundoff errors to be amplified by about two orders of magnitude. So if we are just barely meeting the default 1e-11 Cactus tolerance in the evolved variables, then the constraints might be worse. But I don't really understand this.
However, I have just tried using the Intel compiler options "-fp-model precise -fp-model source" as recommended by Intel if you want reproducibility of floating point computations,
http://software.intel.com/en-us/articles/consistency-of-floating-point-resul...
and if you do that, the test suites pass with "files identical" with 1, 2 and 8 threads on my laptop. So this confirms that the problem comes from a behaviour of the compiler, and not from a problem in the code. Of course we would like to test the compiler settings we actually use in production, and the above settings probably cause a significant loss of optimisation.
That is good to know.
However, we currently cannot do that on a per variable-basis and cannot implement and test this to be ready for the release. Thus, I suggest to split the testsuite for the moment into two almost identical suites, one generating only the contraint output, and one generating all other output. Those two can then have different tolerance levels. Once we implemented something in Cactus which can set the tolerance on a per-file basis, we could combine those again.
Is it right to set the tolerance on what we observe, rather than on what we conclude we should expect? It feels a bit like cheating to me.
Good look with tracing roundoff errors through a BSSN RHS evaluation given some data....
Cheers,
Peter
On Thu, Jun 17, 2010 at 12:35:20AM +0200, Ian Hinder wrote:
and if you do that, the test suites pass with "files identical" with 1, 2 and 8 threads on my laptop. So this confirms that the problem comes from a behaviour of the compiler, and not from a problem in the code. Of course we would like to test the compiler settings we actually use in production, and the above settings probably cause a significant loss of optimisation.
We should investigate how much this really influences performance and I tend towards using the more accurate versions if the runtime does not increase by to much.
Is it right to set the tolerance on what we observe, rather than on what we conclude we should expect? It feels a bit like cheating to me.
That is what I tried to say. I expect differences of that level if there are differences in the metric on a lower level. I didn't talk about single cpu vs. smp, but rather just reordering terms in equations, mathematically correct but numerically different. If this happens to the metric and you loose some digits there due to cancelations, this can only be amplified in the curvature and quantities like the constraints or psi4. This way even roundoff can amplify above the standard 1.e-12 absolute error tolerance.
Frank
users@lists.einsteintoolkit.org