I ran the Simfactory benchmark for ML_BSSN on both the current version and the "rewrite" branch to see whether this branch is ready for production use. I ran this benchmark on a single node of Shelob at LSU. In both cases, using 2 OpenMP threads and 8 MPI processes per node was fastest, so I am reporting these results below. Since I was interested in the performance of McLachlan, this is a unigrid vacuum benchmark using fourth order differencing.
One noteworthy difference is that dissipation as implemented in the "rewrite" branch is finally approximately as fast as thorn Dissipation, and I have thus used this option for the "rewrite" branch.
Here are the high-level results:
current: 3.03136e-06 sec per grid point rewrite: 2.85734e-06 sec per grid point
That is, the rewrite branch is about 5% faster.
More detailed timing results for CCTK_EVOL are as follows:
Current:
----------------------------------------------------------------------------------------
Time Time Imblnc Timer gettimeof getrusage
percent secs percent secs secs
----------------------------------------------------------------------------------------
97.8% 373.1 0.0% |_CallEvol 373.1 738.5
97.8% 373.1 0.0% | |_CCTK_EVOL 373.1 738.4
97.8% 373.0 0.0% | | |_CallFunction 373 738.3
10.3% 39.3 36.9% | | | |_syncs 62.27 122.6
10.3% 39.3 36.9% | | | | |_Sync 62.25 122.6
2.1% 8.0 18.6% | | | | | |_comm_state[3].state_fill_sen 9.839 19.49
6.1% 23.2 46.3% | | | | | |_comm_state[6].state_do_some_ 43.18 84.73
1.5% 5.5 1.8% | | | | | |_comm_state[7].state_empty_re 5.61 10.84
87.3% 333.2 6.4% | | | |_thorns 310.3 614.5
21.8% 83.0 3.2% | | | | |_ML_BSSN_Advect 84.45 168.8
8.6% 32.9 51.3% | | | | |_ML_BSSN_NewRad 0.0354 0.05899
14.1% 53.6 1.9% | | | | |_ML_BSSN_RHS1 54.68 109.1
16.0% 61.1 2.0% | | | | |_ML_BSSN_RHS2 62.35 124.5
5.1% 19.4 9.0% | | | | |_ML_BSSN_convertToADMBase 18.63 37.17
4.2% 16.1 3.1% | | | | |_ML_BSSN_convertToADMBaseDtLaps 16.11 31.98
2.4% 9.1 20.0% | | | | |_ML_BSSN_enforce 11.34 22.65
8.3% 31.7 15.3% | | | | |_MoL_Add 33.12 62.48
1.0% 3.8 48.9% | | | | |_ReflectionSymmetry_Apply 7.461 14.93
5.2% 19.8 3.1% | | | | |_dissipation_add 20.47 40.87
Rewrite:
--------------------------------------------------------------------------------------------------
Time Time Imblnc Timer gettimeof getrusage
percent secs percent secs secs
--------------------------------------------------------------------------------------------------
99.0% 356.1 0.0% |_CallEvol 356.1 703.8
98.9% 356.0 0.0% | |_CCTK_EVOL 356.1 703.8
98.9% 356.0 0.0% | | |_CallFunction 356 703.7
14.0% 50.4 37.3% | | | |_syncs 80.43 158.2
14.0% 50.4 37.3% | | | | |_Sync 80.42 158.2
2.1% 7.7 7.7% | | | | | |_comm_state[3].state_fill_send_buffers. 8.341 16.51
9.7% 35.0 45.2% | | | | | |_comm_state[6].state_do_some_work.step 63.77 125.2
1.5% 5.6 2.5% | | | | | |_comm_state[7].state_empty_recv_buffers 5.629 10.87
84.8% 305.1 9.9% | | | |_thorns 275.2 544.4
5.5% 19.8 7.1% | | | | |_ML_BSSN_ADMBaseEverywhere 19.58 39.15
4.4% 16.0 5.3% | | | | |_ML_BSSN_ADMBaseInterior 15.91 31.63
2.3% 8.3 16.6% | | | | |_ML_BSSN_EnforceEverywhere 10.01 20.02
10.8% 39.0 5.0% | | | | |_ML_BSSN_EvolutionInteriorSplitBy1 38.75 77.28
21.2% 76.4 4.6% | | | | |_ML_BSSN_EvolutionInteriorSplitBy2 75.82 151.5
23.3% 83.9 5.8% | | | | |_ML_BSSN_EvolutionInteriorSplitBy3 83.78 166.8
9.7% 35.0 48.9% | | | | |_ML_BSSN_NewRad 0.1266 0.156
6.0% 21.5 20.7% | | | | |_MoL_Add 23.12 43.16
0.9% 3.3 49.4% | | | | |_ReflectionSymmetry_Apply 6.508 12.97
As can be seen, the non-McLachlan numbers are comparable, although not identical. This is to be expected. The main RHS evaluation is split over several routines in each case; these are (RHS1, RHS2, Advect, dissipation_add) for the current version and (EvolutionInteriorSplitBy[1,2,3]) for the rewrite branch. The way in which the RHS evaluation is actually split is of course different for both cases.
With these numbers in hand, I think we are ready to switch to the rewrite branch.
-erik
On 3 Jul 2015, at 22:38, Erik Schnetter schnetter@cct.lsu.edu wrote:
I ran the Simfactory benchmark for ML_BSSN on both the current version and the "rewrite" branch to see whether this branch is ready for production use. I ran this benchmark on a single node of Shelob at LSU. In both cases, using 2 OpenMP threads and 8 MPI processes per node was fastest, so I am reporting these results below. Since I was interested in the performance of McLachlan, this is a unigrid vacuum benchmark using fourth order differencing.
One noteworthy difference is that dissipation as implemented in the "rewrite" branch is finally approximately as fast as thorn Dissipation, and I have thus used this option for the "rewrite" branch.
Here are the high-level results:
current: 3.03136e-06 sec per grid point rewrite: 2.85734e-06 sec per grid point
That is, the rewrite branch is about 5% faster.
Hi Erik,
That is very reassuring! However, for production use, I would be more interested in 6th or 8th order finite differencing (where the advection stencils become very large), and with Jacobians. If 8th order with Jacobians is at least a similar speed with the rewrite branch, then I would be happy with switching.
On Sat, Jul 4, 2015 at 10:21 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 3 Jul 2015, at 22:38, Erik Schnetter schnetter@cct.lsu.edu wrote:
I ran the Simfactory benchmark for ML_BSSN on both the current version and the "rewrite" branch to see whether this branch is ready for production use. I ran this benchmark on a single node of Shelob at LSU. In both cases, using 2 OpenMP threads and 8 MPI processes per node was fastest, so I am reporting these results below. Since I was interested in the performance of McLachlan, this is a unigrid vacuum benchmark using fourth order differencing.
One noteworthy difference is that dissipation as implemented in the "rewrite" branch is finally approximately as fast as thorn Dissipation, and I have thus used this option for the "rewrite" branch.
Here are the high-level results:
current: 3.03136e-06 sec per grid point rewrite: 2.85734e-06 sec per grid point
That is, the rewrite branch is about 5% faster.
Hi Erik,
That is very reassuring! However, for production use, I would be more interested in 6th or 8th order finite differencing (where the advection stencils become very large), and with Jacobians. If 8th order with Jacobians is at least a similar speed with the rewrite branch, then I would be happy with switching.
Ian
Do you want to suggest a particular benchmark parameter file?
-erik
hi Erik,
You could try the ones at
https://bitbucket.org/ianhinder/cactusbench/src/faea4e13ed4232968e81edd1bbc8...
I haven't updated them in a while, but hopefully the ET is sufficiently backward compatible for them to still work.
I added a second benchmark, using a Thornburg04 patch system, 8th order finite differencing, and 4th order patch interpolation. The results are
original: 8.53935e-06 sec
rewrite: 8.55188e-06 sec
this time with 1 thread per MPI process, since that was most efficient in both cases. Most of the time is spent in inter-patch interpolation, which is much more expensive than in a "regular" case since this benchmark is run on a single node and hence with very small grids.
With these numbers under our belt, can we merge the rewrite branch?
-erik
On Sat, Jul 4, 2015 at 5:19 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
hi Erik,
You could try the ones at
https://bitbucket.org/ianhinder/cactusbench/src/faea4e13ed4232968e81edd1bbc8...
I haven't updated them in a while, but hopefully the ET is sufficiently backward compatible for them to still work.
-- Ian Hinder http://members.aei.mpg.de/ianhin
On 4 Jul 2015, at 17:04, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Sat, Jul 4, 2015 at 10:21 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 3 Jul 2015, at 22:38, Erik Schnetter schnetter@cct.lsu.edu wrote:
I ran the Simfactory benchmark for ML_BSSN on both the current version and the "rewrite" branch to see whether this branch is ready for production use. I ran this benchmark on a single node of Shelob at LSU. In both cases, using 2 OpenMP threads and 8 MPI processes per node was fastest, so I am reporting these results below. Since I was interested in the performance of McLachlan, this is a unigrid vacuum benchmark using fourth order differencing.
One noteworthy difference is that dissipation as implemented in the "rewrite" branch is finally approximately as fast as thorn Dissipation, and I have thus used this option for the "rewrite" branch.
Here are the high-level results:
current: 3.03136e-06 sec per grid point rewrite: 2.85734e-06 sec per grid point
That is, the rewrite branch is about 5% faster.
Hi Erik,
That is very reassuring! However, for production use, I would be more interested in 6th or 8th order finite differencing (where the advection stencils become very large), and with Jacobians. If 8th order with Jacobians is at least a similar speed with the rewrite branch, then I would be happy with switching.
Ian
Do you want to suggest a particular benchmark parameter file?
-erik
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
On 8 Jul 2015, at 15:14, Erik Schnetter schnetter@cct.lsu.edu wrote:
I added a second benchmark, using a Thornburg04 patch system, 8th order finite differencing, and 4th order patch interpolation. The results are
original: 8.53935e-06 sec rewrite: 8.55188e-06 sec
this time with 1 thread per MPI process, since that was most efficient in both cases. Most of the time is spent in inter-patch interpolation, which is much more expensive than in a "regular" case since this benchmark is run on a single node and hence with very small grids.
With these numbers under our belt, can we merge the rewrite branch?
The "jacobian" benchmark that I gave you was still a pure kernel benchmark, involving no interpatch interpolation. It just measured the speed of the RHSs when Jacobians were included. I would also not use a single-threaded benchmark with very small grid sizes; this might have been fastest in this artificial case, but in practice I don't think we would use that configuration. The benchmark you have now run seems to be more of a "complete system" benchmark, which is useful, but different.
I think it is important that the kernel itself has not gotten slower, even if the kernel is not currently a major contributor to runtime. We specifically split out the advection derivatives because they made the code with 8th order and Jacobians a fair bit slower. I would just like to see that this is not still the case with the new version, which has changed the way this is handled.
-erik
On Sat, Jul 4, 2015 at 5:19 PM, Ian Hinder ian.hinder@aei.mpg.de wrote: hi Erik,
You could try the ones at
https://bitbucket.org/ianhinder/cactusbench/src/faea4e13ed4232968e81edd1bbc8...
I haven't updated them in a while, but hopefully the ET is sufficiently backward compatible for them to still work.
-- Ian Hinder http://members.aei.mpg.de/ianhin
On 4 Jul 2015, at 17:04, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Sat, Jul 4, 2015 at 10:21 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 3 Jul 2015, at 22:38, Erik Schnetter schnetter@cct.lsu.edu wrote:
I ran the Simfactory benchmark for ML_BSSN on both the current version and the "rewrite" branch to see whether this branch is ready for production use. I ran this benchmark on a single node of Shelob at LSU. In both cases, using 2 OpenMP threads and 8 MPI processes per node was fastest, so I am reporting these results below. Since I was interested in the performance of McLachlan, this is a unigrid vacuum benchmark using fourth order differencing.
One noteworthy difference is that dissipation as implemented in the "rewrite" branch is finally approximately as fast as thorn Dissipation, and I have thus used this option for the "rewrite" branch.
Here are the high-level results:
current: 3.03136e-06 sec per grid point rewrite: 2.85734e-06 sec per grid point
That is, the rewrite branch is about 5% faster.
Hi Erik,
That is very reassuring! However, for production use, I would be more interested in 6th or 8th order finite differencing (where the advection stencils become very large), and with Jacobians. If 8th order with Jacobians is at least a similar speed with the rewrite branch, then I would be happy with switching.
Ian
Do you want to suggest a particular benchmark parameter file?
-erik
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
On 8 Jul 2015, at 16:53, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 15:14, Erik Schnetter schnetter@cct.lsu.edu wrote:
I added a second benchmark, using a Thornburg04 patch system, 8th order finite differencing, and 4th order patch interpolation. The results are
original: 8.53935e-06 sec rewrite: 8.55188e-06 sec
this time with 1 thread per MPI process, since that was most efficient in both cases. Most of the time is spent in inter-patch interpolation, which is much more expensive than in a "regular" case since this benchmark is run on a single node and hence with very small grids.
With these numbers under our belt, can we merge the rewrite branch?
The "jacobian" benchmark that I gave you was still a pure kernel benchmark, involving no interpatch interpolation. It just measured the speed of the RHSs when Jacobians were included. I would also not use a single-threaded benchmark with very small grid sizes; this might have been fastest in this artificial case, but in practice I don't think we would use that configuration. The benchmark you have now run seems to be more of a "complete system" benchmark, which is useful, but different.
I think it is important that the kernel itself has not gotten slower, even if the kernel is not currently a major contributor to runtime. We specifically split out the advection derivatives because they made the code with 8th order and Jacobians a fair bit slower. I would just like to see that this is not still the case with the new version, which has changed the way this is handled.
I have now run my benchmarks on both the original and the rewritten McLachlan. I seem to find that the ML_BSSN_* functions in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns, excluding the constraint calculations, are between 11% and 15% slower with the rewrite branch, depending on the details of the evolution. See attached plot. This is on Datura with quite old CPUs (Intel Xeon CPU X5650 2.67GHz).
On Fri, Jul 24, 2015 at 11:57 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 16:53, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 15:14, Erik Schnetter schnetter@cct.lsu.edu wrote:
I added a second benchmark, using a Thornburg04 patch system, 8th order finite differencing, and 4th order patch interpolation. The results are
original: 8.53935e-06 sec rewrite: 8.55188e-06 sec
this time with 1 thread per MPI process, since that was most efficient in both cases. Most of the time is spent in inter-patch interpolation, which is much more expensive than in a "regular" case since this benchmark is run on a single node and hence with very small grids.
With these numbers under our belt, can we merge the rewrite branch?
The "jacobian" benchmark that I gave you was still a pure kernel benchmark, involving no interpatch interpolation. It just measured the speed of the RHSs when Jacobians were included. I would also not use a single-threaded benchmark with very small grid sizes; this might have been fastest in this artificial case, but in practice I don't think we would use that configuration. The benchmark you have now run seems to be more of a "complete system" benchmark, which is useful, but different.
I think it is important that the kernel itself has not gotten slower, even if the kernel is not currently a major contributor to runtime. We specifically split out the advection derivatives because they made the code with 8th order and Jacobians a fair bit slower. I would just like to see that this is not still the case with the new version, which has changed the way this is handled.
I have now run my benchmarks on both the original and the rewritten McLachlan. I seem to find that the ML_BSSN_* functions in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns, excluding the constraint calculations, are between 11% and 15% slower with the rewrite branch, depending on the details of the evolution. See attached plot. This is on Datura with quite old CPUs (Intel Xeon CPU X5650 2.67GHz).
What exactly do you measure -- which bins or routines? Does this involve communication? Are you using thorn Dissipation?
-erik
On 24 Jul 2015, at 19:15, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 11:57 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 16:53, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 15:14, Erik Schnetter schnetter@cct.lsu.edu wrote:
I added a second benchmark, using a Thornburg04 patch system, 8th order finite differencing, and 4th order patch interpolation. The results are
original: 8.53935e-06 sec rewrite: 8.55188e-06 sec
this time with 1 thread per MPI process, since that was most efficient in both cases. Most of the time is spent in inter-patch interpolation, which is much more expensive than in a "regular" case since this benchmark is run on a single node and hence with very small grids.
With these numbers under our belt, can we merge the rewrite branch?
The "jacobian" benchmark that I gave you was still a pure kernel benchmark, involving no interpatch interpolation. It just measured the speed of the RHSs when Jacobians were included. I would also not use a single-threaded benchmark with very small grid sizes; this might have been fastest in this artificial case, but in practice I don't think we would use that configuration. The benchmark you have now run seems to be more of a "complete system" benchmark, which is useful, but different.
I think it is important that the kernel itself has not gotten slower, even if the kernel is not currently a major contributor to runtime. We specifically split out the advection derivatives because they made the code with 8th order and Jacobians a fair bit slower. I would just like to see that this is not still the case with the new version, which has changed the way this is handled.
I have now run my benchmarks on both the original and the rewritten McLachlan. I seem to find that the ML_BSSN_* functions in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns, excluding the constraint calculations, are between 11% and 15% slower with the rewrite branch, depending on the details of the evolution. See attached plot. This is on Datura with quite old CPUs (Intel Xeon CPU X5650 2.67GHz).
What exactly do you measure -- which bins or routines? Does this involve communication? Are you using thorn Dissipation?
I take all the timers in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns that start with ML_BSSN_ and eliminate the ones containing "constraints" (case insensitive). This is running on two processes, one node, 6 threads per node. Threads are correctly bound to cores. There is ghostzone exchange between the processes, so yes, there is communication in the ML_BSSN_SelectBCs SYNC calls, but it is node-local.
On Fri, Jul 24, 2015 at 1:39 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 24 Jul 2015, at 19:15, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 11:57 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 16:53, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 15:14, Erik Schnetter schnetter@cct.lsu.edu wrote:
I added a second benchmark, using a Thornburg04 patch system, 8th order finite differencing, and 4th order patch interpolation. The results are
original: 8.53935e-06 sec rewrite: 8.55188e-06 sec
this time with 1 thread per MPI process, since that was most efficient in both cases. Most of the time is spent in inter-patch interpolation, which is much more expensive than in a "regular" case since this benchmark is run on a single node and hence with very small grids.
With these numbers under our belt, can we merge the rewrite branch?
The "jacobian" benchmark that I gave you was still a pure kernel benchmark, involving no interpatch interpolation. It just measured the speed of the RHSs when Jacobians were included. I would also not use a single-threaded benchmark with very small grid sizes; this might have been fastest in this artificial case, but in practice I don't think we would use that configuration. The benchmark you have now run seems to be more of a "complete system" benchmark, which is useful, but different.
I think it is important that the kernel itself has not gotten slower, even if the kernel is not currently a major contributor to runtime. We specifically split out the advection derivatives because they made the code with 8th order and Jacobians a fair bit slower. I would just like to see that this is not still the case with the new version, which has changed the way this is handled.
I have now run my benchmarks on both the original and the rewritten McLachlan. I seem to find that the ML_BSSN_* functions in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns, excluding the constraint calculations, are between 11% and 15% slower with the rewrite branch, depending on the details of the evolution. See attached plot. This is on Datura with quite old CPUs (Intel Xeon CPU X5650 2.67GHz).
What exactly do you measure -- which bins or routines? Does this involve communication? Are you using thorn Dissipation?
I take all the timers in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns that start with ML_BSSN_ and eliminate the ones containing "constraints" (case insensitive). This is running on two processes, one node, 6 threads per node. Threads are correctly bound to cores. There is ghostzone exchange between the processes, so yes, there is communication in the ML_BSSN_SelectBCs SYNC calls, but it is node-local.
Can you include thorn Dissipation in the "before" case, and use McLachlan's dissipation in the "after" case?
-erik
On 24 Jul 2015, at 19:42, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 1:39 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 24 Jul 2015, at 19:15, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 11:57 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 16:53, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 15:14, Erik Schnetter schnetter@cct.lsu.edu wrote:
I added a second benchmark, using a Thornburg04 patch system, 8th order finite differencing, and 4th order patch interpolation. The results are
original: 8.53935e-06 sec rewrite: 8.55188e-06 sec
this time with 1 thread per MPI process, since that was most efficient in both cases. Most of the time is spent in inter-patch interpolation, which is much more expensive than in a "regular" case since this benchmark is run on a single node and hence with very small grids.
With these numbers under our belt, can we merge the rewrite branch?
The "jacobian" benchmark that I gave you was still a pure kernel benchmark, involving no interpatch interpolation. It just measured the speed of the RHSs when Jacobians were included. I would also not use a single-threaded benchmark with very small grid sizes; this might have been fastest in this artificial case, but in practice I don't think we would use that configuration. The benchmark you have now run seems to be more of a "complete system" benchmark, which is useful, but different.
I think it is important that the kernel itself has not gotten slower, even if the kernel is not currently a major contributor to runtime. We specifically split out the advection derivatives because they made the code with 8th order and Jacobians a fair bit slower. I would just like to see that this is not still the case with the new version, which has changed the way this is handled.
I have now run my benchmarks on both the original and the rewritten McLachlan. I seem to find that the ML_BSSN_* functions in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns, excluding the constraint calculations, are between 11% and 15% slower with the rewrite branch, depending on the details of the evolution. See attached plot. This is on Datura with quite old CPUs (Intel Xeon CPU X5650 2.67GHz).
What exactly do you measure -- which bins or routines? Does this involve communication? Are you using thorn Dissipation?
I take all the timers in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns that start with ML_BSSN_ and eliminate the ones containing "constraints" (case insensitive). This is running on two processes, one node, 6 threads per node. Threads are correctly bound to cores. There is ghostzone exchange between the processes, so yes, there is communication in the ML_BSSN_SelectBCs SYNC calls, but it is node-local.
Can you include thorn Dissipation in the "before" case, and use McLachlan's dissipation in the "after" case?
There is no dissipation in either case.
The output data is in
http://git.barrywardell.net/?p=McLachlanBenchmarks.git;h=refs/runs/orig/2015... http://git.barrywardell.net/?p=McLachlanBenchmarks.git;h=refs/runs/rewrite/2...
including the parameter files.
Actually, what I said before was wrong; the timers I am using are under "thorns", not "syncs", so even the node-local communication should not be counted.
On Fri, Jul 24, 2015 at 1:58 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 24 Jul 2015, at 19:42, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 1:39 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 24 Jul 2015, at 19:15, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 11:57 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 16:53, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 15:14, Erik Schnetter schnetter@cct.lsu.edu wrote:
I added a second benchmark, using a Thornburg04 patch system, 8th order finite differencing, and 4th order patch interpolation. The results are
original: 8.53935e-06 sec rewrite: 8.55188e-06 sec
this time with 1 thread per MPI process, since that was most efficient in both cases. Most of the time is spent in inter-patch interpolation, which is much more expensive than in a "regular" case since this benchmark is run on a single node and hence with very small grids.
With these numbers under our belt, can we merge the rewrite branch?
The "jacobian" benchmark that I gave you was still a pure kernel benchmark, involving no interpatch interpolation. It just measured the speed of the RHSs when Jacobians were included. I would also not use a single-threaded benchmark with very small grid sizes; this might have been fastest in this artificial case, but in practice I don't think we would use that configuration. The benchmark you have now run seems to be more of a "complete system" benchmark, which is useful, but different.
I think it is important that the kernel itself has not gotten slower, even if the kernel is not currently a major contributor to runtime. We specifically split out the advection derivatives because they made the code with 8th order and Jacobians a fair bit slower. I would just like to see that this is not still the case with the new version, which has changed the way this is handled.
I have now run my benchmarks on both the original and the rewritten McLachlan. I seem to find that the ML_BSSN_* functions in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns, excluding the constraint calculations, are between 11% and 15% slower with the rewrite branch, depending on the details of the evolution. See attached plot. This is on Datura with quite old CPUs (Intel Xeon CPU X5650 2.67GHz).
What exactly do you measure -- which bins or routines? Does this involve communication? Are you using thorn Dissipation?
I take all the timers in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns that start with ML_BSSN_ and eliminate the ones containing "constraints" (case insensitive). This is running on two processes, one node, 6 threads per node. Threads are correctly bound to cores. There is ghostzone exchange between the processes, so yes, there is communication in the ML_BSSN_SelectBCs SYNC calls, but it is node-local.
Can you include thorn Dissipation in the "before" case, and use McLachlan's dissipation in the "after" case?
There is no dissipation in either case.
The output data is in
http://git.barrywardell.net/?p=McLachlanBenchmarks.git;h=refs/runs/orig/2015...
http://git.barrywardell.net/?p=McLachlanBenchmarks.git;h=refs/runs/rewrite/2...
including the parameter files.
Actually, what I said before was wrong; the timers I am using are under "thorns", not "syncs", so even the node-local communication should not be counted.
McLachlan has not been optimized for runs without dissipation. If you this this is important, then we can introduce a special case. I expect this to improve performance. However, running BSSN without dissipation is not what one would do in production, so I didn't investigate this case.
-erik
On 24 Jul 2015, at 20:39, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 1:58 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 24 Jul 2015, at 19:42, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 1:39 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 24 Jul 2015, at 19:15, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 11:57 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 16:53, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 15:14, Erik Schnetter schnetter@cct.lsu.edu wrote:
I added a second benchmark, using a Thornburg04 patch system, 8th order finite differencing, and 4th order patch interpolation. The results are
original: 8.53935e-06 sec rewrite: 8.55188e-06 sec
this time with 1 thread per MPI process, since that was most efficient in both cases. Most of the time is spent in inter-patch interpolation, which is much more expensive than in a "regular" case since this benchmark is run on a single node and hence with very small grids.
With these numbers under our belt, can we merge the rewrite branch?
The "jacobian" benchmark that I gave you was still a pure kernel benchmark, involving no interpatch interpolation. It just measured the speed of the RHSs when Jacobians were included. I would also not use a single-threaded benchmark with very small grid sizes; this might have been fastest in this artificial case, but in practice I don't think we would use that configuration. The benchmark you have now run seems to be more of a "complete system" benchmark, which is useful, but different.
I think it is important that the kernel itself has not gotten slower, even if the kernel is not currently a major contributor to runtime. We specifically split out the advection derivatives because they made the code with 8th order and Jacobians a fair bit slower. I would just like to see that this is not still the case with the new version, which has changed the way this is handled.
I have now run my benchmarks on both the original and the rewritten McLachlan. I seem to find that the ML_BSSN_* functions in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns, excluding the constraint calculations, are between 11% and 15% slower with the rewrite branch, depending on the details of the evolution. See attached plot. This is on Datura with quite old CPUs (Intel Xeon CPU X5650 2.67GHz).
What exactly do you measure -- which bins or routines? Does this involve communication? Are you using thorn Dissipation?
I take all the timers in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns that start with ML_BSSN_ and eliminate the ones containing "constraints" (case insensitive). This is running on two processes, one node, 6 threads per node. Threads are correctly bound to cores. There is ghostzone exchange between the processes, so yes, there is communication in the ML_BSSN_SelectBCs SYNC calls, but it is node-local.
Can you include thorn Dissipation in the "before" case, and use McLachlan's dissipation in the "after" case?
There is no dissipation in either case.
The output data is in
http://git.barrywardell.net/?p=McLachlanBenchmarks.git;h=refs/runs/orig/2015... http://git.barrywardell.net/?p=McLachlanBenchmarks.git;h=refs/runs/rewrite/2...
including the parameter files.
Actually, what I said before was wrong; the timers I am using are under "thorns", not "syncs", so even the node-local communication should not be counted.
McLachlan has not been optimized for runs without dissipation. If you this this is important, then we can introduce a special case. I expect this to improve performance. However, running BSSN without dissipation is not what one would do in production, so I didn't investigate this case.
I agree that runs without dissipation are not relevant, but since I usually use the Dissipation thorn, I didn't include it in the benchmark, which was a benchmark of McLachlan. I assume that McLachlan now always calculates the dissipation term, even when it is zero, and that is what you mean by "not optimised"? This will introduce a performance regression (if this is the reason for the increased benchmark time, then presumably only on the level of ~15% for the kernel, hence less for a whole simulation) for any simulation which uses dissipation from the Dissipation thorn. Since McLachlan's dissipation was previously very slow, this is presumably what most existing parameter files use.
Regarding switching to use McLachlan for dissipation: McLachlan's dissipation is a bit more limited than the Dissipation thorn; it looks like McLachlan is hard-coded to use dissipation of order 1+fdOrder, rather than the dissipation order being chosen separately. Sometimes lower orders are used as an optimisation (the effect on convergence being judged to be minimal). And actually, critically, there is no way to specify different dissipation orders on different refinement levels. This is typically used in production binary simulations.
Do you think it is faster to use dissipation from McLachlan than to use that provided by Dissipation?
On Fri, Jul 24, 2015 at 3:43 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 24 Jul 2015, at 20:39, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 1:58 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 24 Jul 2015, at 19:42, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 1:39 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 24 Jul 2015, at 19:15, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 11:57 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 16:53, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 15:14, Erik Schnetter schnetter@cct.lsu.edu wrote:
I added a second benchmark, using a Thornburg04 patch system, 8th order finite differencing, and 4th order patch interpolation. The results are
original: 8.53935e-06 sec rewrite: 8.55188e-06 sec
this time with 1 thread per MPI process, since that was most efficient in both cases. Most of the time is spent in inter-patch interpolation, which is much more expensive than in a "regular" case since this benchmark is run on a single node and hence with very small grids.
With these numbers under our belt, can we merge the rewrite branch?
The "jacobian" benchmark that I gave you was still a pure kernel benchmark, involving no interpatch interpolation. It just measured the speed of the RHSs when Jacobians were included. I would also not use a single-threaded benchmark with very small grid sizes; this might have been fastest in this artificial case, but in practice I don't think we would use that configuration. The benchmark you have now run seems to be more of a "complete system" benchmark, which is useful, but different.
I think it is important that the kernel itself has not gotten slower, even if the kernel is not currently a major contributor to runtime. We specifically split out the advection derivatives because they made the code with 8th order and Jacobians a fair bit slower. I would just like to see that this is not still the case with the new version, which has changed the way this is handled.
I have now run my benchmarks on both the original and the rewritten McLachlan. I seem to find that the ML_BSSN_* functions in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns, excluding the constraint calculations, are between 11% and 15% slower with the rewrite branch, depending on the details of the evolution. See attached plot. This is on Datura with quite old CPUs (Intel Xeon CPU X5650 2.67GHz).
What exactly do you measure -- which bins or routines? Does this involve communication? Are you using thorn Dissipation?
I take all the timers in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns that start with ML_BSSN_ and eliminate the ones containing "constraints" (case insensitive). This is running on two processes, one node, 6 threads per node. Threads are correctly bound to cores. There is ghostzone exchange between the processes, so yes, there is communication in the ML_BSSN_SelectBCs SYNC calls, but it is node-local.
Can you include thorn Dissipation in the "before" case, and use McLachlan's dissipation in the "after" case?
There is no dissipation in either case.
The output data is in
http://git.barrywardell.net/?p=McLachlanBenchmarks.git;h=refs/runs/orig/2015...
http://git.barrywardell.net/?p=McLachlanBenchmarks.git;h=refs/runs/rewrite/2...
including the parameter files.
Actually, what I said before was wrong; the timers I am using are under "thorns", not "syncs", so even the node-local communication should not be counted.
McLachlan has not been optimized for runs without dissipation. If you this this is important, then we can introduce a special case. I expect this to improve performance. However, running BSSN without dissipation is not what one would do in production, so I didn't investigate this case.
I agree that runs without dissipation are not relevant, but since I usually use the Dissipation thorn, I didn't include it in the benchmark, which was a benchmark of McLachlan. I assume that McLachlan now always calculates the dissipation term, even when it is zero, and that is what you mean by "not optimised"? This will introduce a performance regression (if this is the reason for the increased benchmark time, then presumably only on the level of ~15% for the kernel, hence less for a whole simulation) for any simulation which uses dissipation from the Dissipation thorn. Since McLachlan's dissipation was previously very slow, this is presumably what most existing parameter files use.
Regarding switching to use McLachlan for dissipation: McLachlan's dissipation is a bit more limited than the Dissipation thorn; it looks like McLachlan is hard-coded to use dissipation of order 1+fdOrder, rather than the dissipation order being chosen separately. Sometimes lower orders are used as an optimisation (the effect on convergence being judged to be minimal). And actually, critically, there is no way to specify different dissipation orders on different refinement levels. This is typically used in production binary simulations.
In other words, you are asking for a version of ML_BSSN where it is efficient to not use dissipation. Currently, that means that dissipation is disabled. The question is -- should this be the default?
Do you think it is faster to use dissipation from McLachlan than to use
that provided by Dissipation?
Yes, I think so.
-erik
On 24 Jul 2015, at 23:01, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 3:43 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 24 Jul 2015, at 20:39, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 1:58 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 24 Jul 2015, at 19:42, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 1:39 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 24 Jul 2015, at 19:15, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 11:57 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 16:53, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 15:14, Erik Schnetter schnetter@cct.lsu.edu wrote:
I added a second benchmark, using a Thornburg04 patch system, 8th order finite differencing, and 4th order patch interpolation. The results are
original: 8.53935e-06 sec rewrite: 8.55188e-06 sec
this time with 1 thread per MPI process, since that was most efficient in both cases. Most of the time is spent in inter-patch interpolation, which is much more expensive than in a "regular" case since this benchmark is run on a single node and hence with very small grids.
With these numbers under our belt, can we merge the rewrite branch?
The "jacobian" benchmark that I gave you was still a pure kernel benchmark, involving no interpatch interpolation. It just measured the speed of the RHSs when Jacobians were included. I would also not use a single-threaded benchmark with very small grid sizes; this might have been fastest in this artificial case, but in practice I don't think we would use that configuration. The benchmark you have now run seems to be more of a "complete system" benchmark, which is useful, but different.
I think it is important that the kernel itself has not gotten slower, even if the kernel is not currently a major contributor to runtime. We specifically split out the advection derivatives because they made the code with 8th order and Jacobians a fair bit slower. I would just like to see that this is not still the case with the new version, which has changed the way this is handled.
I have now run my benchmarks on both the original and the rewritten McLachlan. I seem to find that the ML_BSSN_* functions in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns, excluding the constraint calculations, are between 11% and 15% slower with the rewrite branch, depending on the details of the evolution. See attached plot. This is on Datura with quite old CPUs (Intel Xeon CPU X5650 2.67GHz).
What exactly do you measure -- which bins or routines? Does this involve communication? Are you using thorn Dissipation?
I take all the timers in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns that start with ML_BSSN_ and eliminate the ones containing "constraints" (case insensitive). This is running on two processes, one node, 6 threads per node. Threads are correctly bound to cores. There is ghostzone exchange between the processes, so yes, there is communication in the ML_BSSN_SelectBCs SYNC calls, but it is node-local.
Can you include thorn Dissipation in the "before" case, and use McLachlan's dissipation in the "after" case?
There is no dissipation in either case.
The output data is in
http://git.barrywardell.net/?p=McLachlanBenchmarks.git;h=refs/runs/orig/2015... http://git.barrywardell.net/?p=McLachlanBenchmarks.git;h=refs/runs/rewrite/2...
including the parameter files.
Actually, what I said before was wrong; the timers I am using are under "thorns", not "syncs", so even the node-local communication should not be counted.
McLachlan has not been optimized for runs without dissipation. If you this this is important, then we can introduce a special case. I expect this to improve performance. However, running BSSN without dissipation is not what one would do in production, so I didn't investigate this case.
I agree that runs without dissipation are not relevant, but since I usually use the Dissipation thorn, I didn't include it in the benchmark, which was a benchmark of McLachlan. I assume that McLachlan now always calculates the dissipation term, even when it is zero, and that is what you mean by "not optimised"? This will introduce a performance regression (if this is the reason for the increased benchmark time, then presumably only on the level of ~15% for the kernel, hence less for a whole simulation) for any simulation which uses dissipation from the Dissipation thorn. Since McLachlan's dissipation was previously very slow, this is presumably what most existing parameter files use.
Regarding switching to use McLachlan for dissipation: McLachlan's dissipation is a bit more limited than the Dissipation thorn; it looks like McLachlan is hard-coded to use dissipation of order 1+fdOrder, rather than the dissipation order being chosen separately. Sometimes lower orders are used as an optimisation (the effect on convergence being judged to be minimal). And actually, critically, there is no way to specify different dissipation orders on different refinement levels. This is typically used in production binary simulations.
In other words, you are asking for a version of ML_BSSN where it is efficient to not use dissipation. Currently, that means that dissipation is disabled. The question is -- should this be the default?
Do you think it is faster to use dissipation from McLachlan than to use that provided by Dissipation?
Yes, I think so.
I don't know. Without knowing performance numbers, it is difficult to judge. Since people may be using McLachlan's dissipation in their parameter files (even though it is slow), it's probably not a good idea to disable it by default.
Is it possible to make McLachlan efficient when dissipation is disabled, but keep the code for it there? e.g. by wrapping it in a conditional? If the condition is a scalar, this should be fine even with vectorisation, no?
If I remember correctly, the default is zero dissipation, so by default it already is effectively off now.
Frank
On 25 Jul 2015, at 10:35, Frank Loeffler knarf@cct.lsu.edu wrote:
If I remember correctly, the default is zero dissipation, so by default it already is effectively off now.
I'm referring to parameter files which set the McLachlan dissipation parameter. Those will fail if we disable dissipation (at Kranc time).
Am 25. Juli 2015 11:35:27 MESZ, schrieb Ian Hinder ian.hinder@aei.mpg.de:
On 25 Jul 2015, at 10:35, Frank Loeffler knarf@cct.lsu.edu wrote:
If I remember correctly, the default is zero dissipation, so by
default it already is effectively off now.
I'm referring to parameter files which set the McLachlan dissipation parameter. Those will fail if we disable dissipation (at Kranc time).
I didn't think you mean at Kranc time. I would think we shouldn't do that. Rerunning Kranc shouldn't be necessary to enable something like dissipation.
Frank
On 25 Jul 2015, at 12:26, Frank Loeffler knarf@cct.lsu.edu wrote:
Am 25. Juli 2015 11:35:27 MESZ, schrieb Ian Hinder ian.hinder@aei.mpg.de:
On 25 Jul 2015, at 10:35, Frank Loeffler knarf@cct.lsu.edu wrote:
If I remember correctly, the default is zero dissipation, so by
default it already is effectively off now.
I'm referring to parameter files which set the McLachlan dissipation parameter. Those will fail if we disable dissipation (at Kranc time).
I didn't think you mean at Kranc time. I would think we shouldn't do that. Rerunning Kranc shouldn't be necessary to enable something like dissipation.
Actually I do. On the rewrite branch, McLachlan has a switch to enable dissipation at Kranc time. Once enabled, it is unconditionally computed at run time, which takes a certain amount of time. This cannot currently be disabled at run time. Erik was asking (I think) whether this Kranc-time dissipation should be disabled, since in the case that users want to use dissipation from the Dissipation thorn, they incur an unavoidable performance penalty (dissipation terms will be computed twice, though one of them will be multiplied by zero).
Dissipation can be provided:
1. By the Dissipation thorn, which provides several useful (needed) features, such as being able to control the dissipation strength on a given refinement level, the dissipation order, as well as having some options for excision masks (which I have never used, but others might); or 2. By McLachlan, which is more limited, but which Erik expects to be faster
The pre-rewrite version of McLachlan had very slow dissipation, but it was not executed if the dissipation strength was set to 0 at runtime, so this problem did not arise.
On Sat, Jul 25, 2015 at 5:46 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 25 Jul 2015, at 12:26, Frank Loeffler knarf@cct.lsu.edu wrote:
Am 25. Juli 2015 11:35:27 MESZ, schrieb Ian Hinder <ian.hinder@aei.mpg.de
:
On 25 Jul 2015, at 10:35, Frank Loeffler knarf@cct.lsu.edu wrote:
If I remember correctly, the default is zero dissipation, so by
default it already is effectively off now.
I'm referring to parameter files which set the McLachlan dissipation parameter. Those will fail if we disable dissipation (at Kranc time).
I didn't think you mean at Kranc time. I would think we shouldn't do that. Rerunning Kranc shouldn't be necessary to enable something like dissipation.
Actually I do. On the rewrite branch, McLachlan has a switch to enable dissipation at Kranc time. Once enabled, it is unconditionally computed at run time, which takes a certain amount of time. This cannot currently be disabled at run time. Erik was asking (I think) whether this Kranc-time dissipation should be disabled, since in the case that users want to use dissipation from the Dissipation thorn, they incur an unavoidable performance penalty (dissipation terms will be computed twice, though one of them will be multiplied by zero).
To implement dissipation efficiently, it is calculated at the same time as other derivatives. This avoids having multiple loops over the grid functions. Previously, McLachlan applied the dissipation in a routine of its own that could easily be scheduled or not. Now, it is applied when the RHS is calculated, and Kranc does not offer a mechanism that would avoid calculating the dissipation stencils based on a parameter. It would easily be possible to add a parameter that skips adding dissipation to the RHS, but this is essentially the same as setting the dissipation strength to zero. However, even if it is not added, the stencils is still calculated. Hence this decision needs to be made at Kranc time, until Kranc is extended to offer this functionality.
In principle, calculating the dissipation terms while the RHS is calculated should be faster. This also seems to be the case currently -- old ML without dissipation + extra dissipation is slightly slower than new ML with dissipation. However, this is not the case which interests Ian; Ian wants to use an external dissipation mechanism. Hence I now created ML_BSSN_ND (No Dissipation) that can be used to test this. I don't think this should be the default; rather, we should ask people to give this a try and measure performance, and we should also add missing features if necessary.
-erik
Dissipation can be provided:
- By the Dissipation thorn, which provides several useful (needed)
features, such as being able to control the dissipation strength on a given refinement level, the dissipation order, as well as having some options for excision masks (which I have never used, but others might); or 2. By McLachlan, which is more limited, but which Erik expects to be faster
The pre-rewrite version of McLachlan had very slow dissipation, but it was not executed if the dissipation strength was set to 0 at runtime, so this problem did not arise.
-- Ian Hinder http://members.aei.mpg.de/ianhin
On Fri, Jul 24, 2015 at 10:57 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 16:53, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 15:14, Erik Schnetter schnetter@cct.lsu.edu wrote:
I added a second benchmark, using a Thornburg04 patch system, 8th order finite differencing, and 4th order patch interpolation. The results are
original: 8.53935e-06 sec rewrite: 8.55188e-06 sec
this time with 1 thread per MPI process, since that was most efficient in both cases. Most of the time is spent in inter-patch interpolation, which is much more expensive than in a "regular" case since this benchmark is run on a single node and hence with very small grids.
With these numbers under our belt, can we merge the rewrite branch?
The "jacobian" benchmark that I gave you was still a pure kernel benchmark, involving no interpatch interpolation. It just measured the speed of the RHSs when Jacobians were included. I would also not use a single-threaded benchmark with very small grid sizes; this might have been fastest in this artificial case, but in practice I don't think we would use that configuration. The benchmark you have now run seems to be more of a "complete system" benchmark, which is useful, but different.
I think it is important that the kernel itself has not gotten slower, even if the kernel is not currently a major contributor to runtime. We specifically split out the advection derivatives because they made the code with 8th order and Jacobians a fair bit slower. I would just like to see that this is not still the case with the new version, which has changed the way this is handled.
I have now run my benchmarks on both the original and the rewritten McLachlan. I seem to find that the ML_BSSN_* functions in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns, excluding the constraint calculations, are between 11% and 15% slower with the rewrite branch, depending on the details of the evolution. See attached plot. This is on Datura with quite old CPUs (Intel Xeon CPU X5650 2.67GHz).
I just realized that you probably used the wrong rhs_evaluation method for McLachlan. While improving performance, I implemented 3 different ways to evaluate the RHS: (1) all in one routine, (2) split manually, and (3) split semi-automatically by Kranc. (2) and (3) are identical for practical purposes, and thus (2) should not be used. In my benchmarks, I always explicitly specified (3). However, the default in McLachlan is still at (1), and thus likely not as efficient as it should be.
The parameter setting ML_BSSN::rhs_evaluation = "splitBy" chooses (3).
I will soon push McLachlan changes to make (3) the default and to remove (2).
-erik
On 29 Jul 2015, at 18:15, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Jul 24, 2015 at 10:57 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 16:53, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 8 Jul 2015, at 15:14, Erik Schnetter schnetter@cct.lsu.edu wrote:
I added a second benchmark, using a Thornburg04 patch system, 8th order finite differencing, and 4th order patch interpolation. The results are
original: 8.53935e-06 sec rewrite: 8.55188e-06 sec
this time with 1 thread per MPI process, since that was most efficient in both cases. Most of the time is spent in inter-patch interpolation, which is much more expensive than in a "regular" case since this benchmark is run on a single node and hence with very small grids.
With these numbers under our belt, can we merge the rewrite branch?
The "jacobian" benchmark that I gave you was still a pure kernel benchmark, involving no interpatch interpolation. It just measured the speed of the RHSs when Jacobians were included. I would also not use a single-threaded benchmark with very small grid sizes; this might have been fastest in this artificial case, but in practice I don't think we would use that configuration. The benchmark you have now run seems to be more of a "complete system" benchmark, which is useful, but different.
I think it is important that the kernel itself has not gotten slower, even if the kernel is not currently a major contributor to runtime. We specifically split out the advection derivatives because they made the code with 8th order and Jacobians a fair bit slower. I would just like to see that this is not still the case with the new version, which has changed the way this is handled.
I have now run my benchmarks on both the original and the rewritten McLachlan. I seem to find that the ML_BSSN_* functions in Evolve/CallEvol/CCTK_EVOL/CallFunction/thorns, excluding the constraint calculations, are between 11% and 15% slower with the rewrite branch, depending on the details of the evolution. See attached plot. This is on Datura with quite old CPUs (Intel Xeon CPU X5650 2.67GHz).
I just realized that you probably used the wrong rhs_evaluation method for McLachlan. While improving performance, I implemented 3 different ways to evaluate the RHS: (1) all in one routine, (2) split manually, and (3) split semi-automatically by Kranc. (2) and (3) are identical for practical purposes, and thus (2) should not be used. In my benchmarks, I always explicitly specified (3). However, the default in McLachlan is still at (1), and thus likely not as efficient as it should be.
The parameter setting ML_BSSN::rhs_evaluation = "splitBy" chooses (3).
I will soon push McLachlan changes to make (3) the default and to remove (2).
Hi Erik,
Regarding the dissipation discussion: would it be possible to select whether to evaluate the dissipation terms in McLachlan using a runtime parameter; i.e. by changing
Dissipation[var_] := IfDiss[epsdiss[ux] PDdiss[var, lx], 0];
to
Dissipation[var_] := IfDiss[IfThen[useDissipation, epsdiss[ux] PDdiss[var, lx], 0], 0];
We can choose a better name for useDissipation. This would be constant across the loop, so the compiler should still be able to vectorise.
users@lists.einsteintoolkit.org