Hi, I did some strong scaling tests on Stampede (https://www.xsede.org/tacc-stampede) with both Whisky and GRHydro (using the development version of ET in both cases). I used a Carpet par file that Roberto DePietri provided me and that he used for similar tests on an Italian cluster (I have attached the GRHydro version, the Whisky one is similar except for using Whisky instead of GRHydro). I used both Intel MPI and Mvapich and I did both pure MPI and MPI/OpenMP runs.
I have attached a text file with my results. The first column is the name of the run (if it starts with mvapich it used mvapich otherwise it used Intel MPI), the second one is the number of cores (option --procs in simfactory), the third one the number of threads (--num-threads), the fourth one the time in seconds spent in CCTK_EVOL, and the fifth one the walltime in seconds (i.e., the total time used by the run as measured on the cluster). I have also attached a couple of figures that show CCTK_EVOL vs #cores and walltime vs #cores (only for Intel MPI runs).
First of all, in pure MPI runs (--num-threads=1) I was unable to run on more than 1024 cores using Intel MPI (the run was just crashing before iteration zero or hanging up). No problem instead when using --num-threads=8 or --num-threads=16. I also noticed that scaling was particularly bad in pure MPI runs and that a lot of time was spent outside CCTK_EVOL (both with Intel MPI and MVAPICH). After speaking with Roberto, I found out that the problem is due to 1D ASCII output (which is active in that parfile) and that makes the runs particularly slow above ~100 cores on this machine. In plot_scaling_walltime_all.pdf I plot also two pure MPI runs, but without 1D ASCII output and the scaling is much better in this case (the time spent in CCTK_EVOL is identical to the case with 1D output and hence I didn't plot them in the other figure). I didn't try using 1D hdf5 output instead, does anyone use it?
According to my tests, --num-threads=16 performs better than --num-threads=8 (which is the current default value in simfactory) and Intel MPI seems to be better than MVAPICH. Is there a particular reason for using 8 instead of 16 threads as the default simfactory value on Stampede?
Let me know if you have any comment or suggestion.
Cheers, Bruno
Dr. Bruno Giacomazzo JILA - University of Colorado 440 UCB Boulder, CO 80309 USA
Tel. : +1 303-492-5170 Fax : +1 303-492-5235 email : bruno.giacomazzo@jila.colorado.edu web: http://www.brunogiacomazzo.org
---------------------------------------------------------------------- There are only 10 types of people in the world: Those who understand binary, and those who don't ----------------------------------------------------------------------
Hi,
On Thu, Feb 21, 2013 at 06:31:23PM -0700, Bruno Giacomazzo wrote:
I found out that the problem is due to 1D ASCII output (which is active in that parfile) and that makes the runs particularly slow above ~100 cores on this machine.
That is not not just on this machine. xD output with x>=1 doesn't scale. Anywhere. Use HDF5, except for debugging where ASCII might be more convenient and you don't care about performance. There isn't much you could do about that without changing CarpetIOASCII.
and hence I didn't plot them in the other figure). I didn't try using 1D hdf5 output instead, does anyone use it?
Not 1D, but 2D, and that works very well.
Frank
Frank,
On Feb 21, 2013, at 8:05 PM, Frank Loeffler wrote:
On Thu, Feb 21, 2013 at 06:31:23PM -0700, Bruno Giacomazzo wrote:
I found out that the problem is due to 1D ASCII output (which is active in that parfile) and that makes the runs particularly slow above ~100 cores on this machine.
That is not not just on this machine. xD output with x>=1 doesn't scale. Anywhere. Use HDF5, except for debugging where ASCII might be more convenient and you don't care about performance. There isn't much you could do about that without changing CarpetIOASCII.
thank you for your comment. It's the first time that I see such a strong impact on runs using only ~100 cores. One of the reasons I did these tests was also that I was experiencing some problems with one of my production runs that was slower on Stampede than on Ranger.
and hence I didn't plot them in the other figure). I didn't try using 1D hdf5 output instead, does anyone use it?
Not 1D, but 2D, and that works very well.
I use 2D hdf5, but sometimes it's useful to have also 1D data (without having to postprocess 2D or 3D hdf5 files). I never used the 1D hdf5 output available in Carpet and I was wondering if other people use it.
Cheers, Bruno
Dr. Bruno Giacomazzo JILA - University of Colorado 440 UCB Boulder, CO 80309 USA
Tel. : +1 303-492-5170 Fax : +1 303-492-5235 email : bruno.giacomazzo@jila.colorado.edu web: http://www.brunogiacomazzo.org
---------------------------------------------------------------------- There are only 10 types of people in the world: Those who understand binary, and those who don't ----------------------------------------------------------------------
Hello all,
and hence I didn't plot them in the other figure). I didn't try using 1D hdf5 output instead, does anyone use it?
Not 1D, but 2D, and that works very well.
So CarpetIOHDF5's 1d output would be expected to be faster than CarpetIOASCII's? Interesting, I had thought them both to be very similar as far as sending around data for proc0 to write was concerned. If that is the case then it is the actual IO routines in the thorns that behave so differently.
Yours, Roland
On Thu, Feb 21, 2013 at 07:52:06PM -0800, Roland Haas wrote:
So CarpetIOHDF5's 1d output would be expected to be faster than CarpetIOASCII's? Interesting, I had thought them both to be very similar as far as sending around data for proc0 to write was concerned. If that is the case then it is the actual IO routines in the thorns that behave so differently.
I was under the impression that CarpetIOASCII does not send any data around. Each process writes to the file, one after the other.
Frank
Hello all,
On Thu, Feb 21, 2013 at 07:52:06PM -0800, Roland Haas wrote:
So CarpetIOHDF5's 1d output would be expected to be faster than CarpetIOASCII's? Interesting, I had thought them both to be very similar as far as sending around data for proc0 to write was concerned. If that is the case then it is the actual IO routines in the thorns that behave so differently.
I was under the impression that CarpetIOASCII does not send any data around. Each process writes to the file, one after the other.
That would be truly horrible I agree. Looking at the code, it does not seem to do so anymore. At least there are instances of "proc == ioproc" and CarpetLib's copy_from.
Yours, Roland
On 22 Feb 2013, at 02:31, Bruno Giacomazzo bruno.giacomazzo@jila.colorado.edu wrote:
Hi, I did some strong scaling tests on Stampede (https://www.xsede.org/tacc-stampede) with both Whisky and GRHydro (using the development version of ET in both cases). I used a Carpet par file that Roberto DePietri provided me and that he used for similar tests on an Italian cluster (I have attached the GRHydro version, the Whisky one is similar except for using Whisky instead of GRHydro). I used both Intel MPI and Mvapich and I did both pure MPI and MPI/OpenMP runs.
I have attached a text file with my results. The first column is the name of the run (if it starts with mvapich it used mvapich otherwise it used Intel MPI), the second one is the number of cores (option --procs in simfactory), the third one the number of threads (--num-threads), the fourth one the time in seconds spent in CCTK_EVOL, and the fifth one the walltime in seconds (i.e., the total time used by the run as measured on the cluster). I have also attached a couple of figures that show CCTK_EVOL vs #cores and walltime vs #cores (only for Intel MPI runs).
First of all, in pure MPI runs (--num-threads=1) I was unable to run on more than 1024 cores using Intel MPI (the run was just crashing before iteration zero or hanging up).
Could it have been running out of memory? Alternatively, some clusters require you to set special job options before running jobs with very large numbers of processes, or to use mpirun in a special way. The Stampede user guide should have details on this. Having said that, no one should be using pure MPI these days anyway.
No problem instead when using --num-threads=8 or --num-threads=16. I also noticed that scaling was particularly bad in pure MPI runs and that a lot of time was spent outside CCTK_EVOL (both with Intel MPI and MVAPICH). After speaking with Roberto, I found out that the problem is due to 1D ASCII output (which is active in that parfile) and that makes the runs particularly slow above ~100 cores on this machine.
Yes, this makes sense. Each process sends all the data to be output to a single process, which then writes the file. So the number of processes involved in the communication will be much larger (by a factor of 6 or 12) if you are using pure MPI. Since it is 1D, you are probably dominated by the number of communicating processes, and associated latency, rather than the actual volume of data being transferred.
In plot_scaling_walltime_all.pdf I plot also two pure MPI runs, but without 1D ASCII output and the scaling is much better in this case (the time spent in CCTK_EVOL is identical to the case with 1D output and hence I didn't plot them in the other figure). I didn't try using 1D hdf5 output instead, does anyone use it?
Yes, I use 1D, 2D and 3D HDF5 output, and it works very well. I only enable ASCII output for testsuites and debugging. HDF5 output is vastly superior: the size of the data is much smaller due to being binary and avoiding extra columns, HDF5 compression can be enabled which reduces it further, and the main reason: HDF5 output is seekable, so I can access iteration 1024 without having to first read and parse iterations 0 to 1023. I don't need a separate "post-processing" phase; I read from the simulation output files directly.
Regarding the scaling issue you mentioned above, 1D HDF5 and 1D ASCII output use the same method for communicating the data to a single process, so I don't expect HDF5 to improve the scaling. The amount of data transferred and the number of processes involved will be the same. The only thing that will improve is the size of the data on the disk, and hence the time taken to write it.
According to my tests, --num-threads=16 performs better than --num-threads=8 (which is the current default value in simfactory) and Intel MPI seems to be better than MVAPICH. Is there a particular reason for using 8 instead of 16 threads as the default simfactory value on Stampede?
Stampede has two processors per node, each with 8 cores. They say "The memory subsystem has 4 channels from each processor's memory controller to 4 DDR3 ECC DIMMS,", which suggests to me that the memory is not associated with each processor (uniform memory architecture; UMA). The usual reason for using one process per processor is that each processor has fast access to only half the memory (nonuniform memory architecture; NUMA). See http://cactuscode.org/pipermail/users/2013-February/thread.html#3295 for a discussion of the situation on SuperMUC. Now, I read that Stampede and SuperMUC are using exactly the same processor, and on SuperMUC, apparently it is NUMA. However, the processors might be deployed in a different way on Stampede. The best approach to investigate this is to enable the thorn hwloc (see that thread) and see what it reports on startup; it should say if the architecture is UMA or NUMA.
I would be careful with interpreting benchmark results. To answer this question definitively, you should run on a single node and look at the timers (all processes) for a single routine which doesn't do any communication; e.g. ML_BSSN_RHS2; for a small number (e.g. 10) of iterations. The processor decomposition and load balancing changes when changing the number of threads might have more of an impact than the memory access speed, and might be different from one grid setup to the next. Similarly, the relation of the local grid size and shape to the cache characteristics of the core might also have a big impact.
Another thing to watch out for: it is essential for good performance to ensure that each thread is bound to a single CPU core ("CPU affinity"). Thorn hwloc will do this automatically; if you don't use it, you rely on the system, or the run scripts, to do this for you, and YMMV.
Bruno
The amount of data output varies greatly between simulations, as well as the intervals between output. And usually, I/O scales very differently than the computation itself. Therefore, it is customary (at least while trying to understand results) to test evolution and output separately. That is, the evolution benchmark would probably output the maximum of rho only, and an I/O benchmark would output (and/or recover) just Minkowski data without time evolution.
Regarding the number of threads: To determine the ideal number of threads to use, we run single-node benchmarks of code that is well parallellised via OpenMP. Code that is not really parallel (e.g. some initial data routines) would distort these results, as would using a large number of cores, since this involves also parallel scalability. Using this "ideal" number of cores, we then run scalability benchmarks to see where parallel scaling breaks down.
For any given physics situation, one then has to strike a compromise: - non-parallel sections of the code prefer using 1 OpenMP thread - parallel MPI scaling prefers using as few MPI processes as possible, i.e. using many OpenMP threads The balance thus shifts with the number of cores -- the more cores you use, the more OpenMP threads you will also want to use to counter-act MPI scaling problems.
As others have said, binding threads to cores and binding memory to processes is also very important, and is visible in particular on single-node benchmarks.
For any benchmark I run, I also look at detailed timer output. The standard Cactus timers are not good enough for this; you will have to use TimerReport or Carpet's timers for this. One interesting quantity is to see what fraction of the time is spent in the actual evolution thorns (not just CCTK_EVOL; this also measures some infrastructure tasks). The other interesting effect to watch is how this time distribution changes as the number of MPI processes increases -- this shows the effect of scaling problems. The latter should e.g. show that ASCII output becomes more time consuming, or that synchronisation or load balancing takes more time.
Finally, I found it very difficult to come up with a "good" benchmark parameter file. Such a benchmark should run both on few and on many cores, should contains all the relevant thorns, should not do I/O, should not contain anything that is known not to scale, should not encounter nans or con2prim problems, should be "close" to actual parameter files that people actually want to use, etc.
I think it's time to create a wiki page for benchmarking! There we could describe (a) the tools available (timers, etc.), (b) the pitfalls to avoid (e.g. measure evolution and I/O separately), and (c) discuss results that we find.
In this case -- I think this means we should improve ASCII output! Writing ASCII files is always slow, but collecting data onto a single process shouldn't be. That's no more than a reduction operation, and with InfiniBand bandwidths of tens of Gigabytes per second, we are very far away from what the hardware performance allows us to do. Let's open a bug report for this. We should either correct this, or should people prominently warn about this.
-erik
On Thu, Feb 21, 2013 at 8:31 PM, Bruno Giacomazzo < bruno.giacomazzo@jila.colorado.edu> wrote:
Hi, I did some strong scaling tests on Stampede ( https://www.xsede.org/tacc-stampede) with both Whisky and GRHydro (using the development version of ET in both cases). I used a Carpet par file that Roberto DePietri provided me and that he used for similar tests on an Italian cluster (I have attached the GRHydro version, the Whisky one is similar except for using Whisky instead of GRHydro). I used both Intel MPI and Mvapich and I did both pure MPI and MPI/OpenMP runs.
I have attached a text file with my results. The first column isthe name of the run (if it starts with mvapich it used mvapich otherwise it used Intel MPI), the second one is the number of cores (option --procs in simfactory), the third one the number of threads (--num-threads), the fourth one the time in seconds spent in CCTK_EVOL, and the fifth one the walltime in seconds (i.e., the total time used by the run as measured on the cluster). I have also attached a couple of figures that show CCTK_EVOL vs #cores and walltime vs #cores (only for Intel MPI runs).
First of all, in pure MPI runs (--num-threads=1) I was unable torun on more than 1024 cores using Intel MPI (the run was just crashing before iteration zero or hanging up). No problem instead when using --num-threads=8 or --num-threads=16. I also noticed that scaling was particularly bad in pure MPI runs and that a lot of time was spent outside CCTK_EVOL (both with Intel MPI and MVAPICH). After speaking with Roberto, I found out that the problem is due to 1D ASCII output (which is active in that parfile) and that makes the runs particularly slow above ~100 cores on this machine. In plot_scaling_walltime_all.pdf I plot also two pure MPI runs, but without 1D ASCII output and the scaling is much better in this case (the time spent in CCTK_EVOL is identical to the case with 1D output and hence I didn't plot them in the other figure). I didn't try using 1D hdf5 output instead, does anyone use it?
According to my tests, --num-threads=16 performs better than--num-threads=8 (which is the current default value in simfactory) and Intel MPI seems to be better than MVAPICH. Is there a particular reason for using 8 instead of 16 threads as the default simfactory value on Stampede?
Let me know if you have any comment or suggestion.Cheers, Bruno
Dr. Bruno Giacomazzo JILA - University of Colorado 440 UCB Boulder, CO 80309 USA
Tel. : +1 303-492-5170 Fax : +1 303-492-5235 email : bruno.giacomazzo@jila.colorado.edu web: http://www.brunogiacomazzo.org
There are only 10 types of people in the world: Those who understand binary, and those who don't
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Dear Erik, Frank, Ian, and Roland, thank you all for the very useful feedback.
Erik, I think it would be nice to have a wiki about benchmarking. It would be especially useful to have a set of "standard" tests (parfiles) that one could use to easily compare ET performance on different clusters.
Ian, thanks for the info, I will try using 1D HDF5 instead of 1D ASCII. In some simulations it's useful to have 1D output ready to use (for the others I will just switch off 1D output)
Cheers, Bruno
On Feb 22, 2013, at 8:09 AM, Erik Schnetter wrote:
Bruno
The amount of data output varies greatly between simulations, as well as the intervals between output. And usually, I/O scales very differently than the computation itself. Therefore, it is customary (at least while trying to understand results) to test evolution and output separately. That is, the evolution benchmark would probably output the maximum of rho only, and an I/O benchmark would output (and/or recover) just Minkowski data without time evolution.
Regarding the number of threads: To determine the ideal number of threads to use, we run single-node benchmarks of code that is well parallellised via OpenMP. Code that is not really parallel (e.g. some initial data routines) would distort these results, as would using a large number of cores, since this involves also parallel scalability. Using this "ideal" number of cores, we then run scalability benchmarks to see where parallel scaling breaks down.
For any given physics situation, one then has to strike a compromise:
- non-parallel sections of the code prefer using 1 OpenMP thread
- parallel MPI scaling prefers using as few MPI processes as possible, i.e. using many OpenMP threads
The balance thus shifts with the number of cores -- the more cores you use, the more OpenMP threads you will also want to use to counter-act MPI scaling problems.
As others have said, binding threads to cores and binding memory to processes is also very important, and is visible in particular on single-node benchmarks.
For any benchmark I run, I also look at detailed timer output. The standard Cactus timers are not good enough for this; you will have to use TimerReport or Carpet's timers for this. One interesting quantity is to see what fraction of the time is spent in the actual evolution thorns (not just CCTK_EVOL; this also measures some infrastructure tasks). The other interesting effect to watch is how this time distribution changes as the number of MPI processes increases -- this shows the effect of scaling problems. The latter should e.g. show that ASCII output becomes more time consuming, or that synchronisation or load balancing takes more time.
Finally, I found it very difficult to come up with a "good" benchmark parameter file. Such a benchmark should run both on few and on many cores, should contains all the relevant thorns, should not do I/O, should not contain anything that is known not to scale, should not encounter nans or con2prim problems, should be "close" to actual parameter files that people actually want to use, etc.
I think it's time to create a wiki page for benchmarking! There we could describe (a) the tools available (timers, etc.), (b) the pitfalls to avoid (e.g. measure evolution and I/O separately), and (c) discuss results that we find.
In this case -- I think this means we should improve ASCII output! Writing ASCII files is always slow, but collecting data onto a single process shouldn't be. That's no more than a reduction operation, and with InfiniBand bandwidths of tens of Gigabytes per second, we are very far away from what the hardware performance allows us to do. Let's open a bug report for this. We should either correct this, or should people prominently warn about this.
-erik
On Thu, Feb 21, 2013 at 8:31 PM, Bruno Giacomazzo bruno.giacomazzo@jila.colorado.edu wrote: Hi, I did some strong scaling tests on Stampede (https://www.xsede.org/tacc-stampede) with both Whisky and GRHydro (using the development version of ET in both cases). I used a Carpet par file that Roberto DePietri provided me and that he used for similar tests on an Italian cluster (I have attached the GRHydro version, the Whisky one is similar except for using Whisky instead of GRHydro). I used both Intel MPI and Mvapich and I did both pure MPI and MPI/OpenMP runs.
I have attached a text file with my results. The first column is the name of the run (if it starts with mvapich it used mvapich otherwise it used Intel MPI), the second one is the number of cores (option --procs in simfactory), the third one the number of threads (--num-threads), the fourth one the time in seconds spent in CCTK_EVOL, and the fifth one the walltime in seconds (i.e., the total time used by the run as measured on the cluster). I have also attached a couple of figures that show CCTK_EVOL vs #cores and walltime vs #cores (only for Intel MPI runs). First of all, in pure MPI runs (--num-threads=1) I was unable to run on more than 1024 cores using Intel MPI (the run was just crashing before iteration zero or hanging up). No problem instead when using --num-threads=8 or --num-threads=16. I also noticed that scaling was particularly bad in pure MPI runs and that a lot of time was spent outside CCTK_EVOL (both with Intel MPI and MVAPICH). After speaking with Roberto, I found out that the problem is due to 1D ASCII output (which is active in that parfile) and that makes the runs particularly slow above ~100 cores on this machine. In plot_scaling_walltime_all.pdf I plot also two pure MPI runs, but without 1D ASCII output and the scaling is much better in this case (the time spent in CCTK_EVOL is identical to the case with 1D output and hence I didn't plot them in the other figure). I didn't try using 1D hdf5 output instead, does anyone use it? According to my tests, --num-threads=16 performs better than --num-threads=8 (which is the current default value in simfactory) and Intel MPI seems to be better than MVAPICH. Is there a particular reason for using 8 instead of 16 threads as the default simfactory value on Stampede? Let me know if you have any comment or suggestion.Cheers, Bruno
Dr. Bruno Giacomazzo JILA - University of Colorado 440 UCB Boulder, CO 80309 USA
Tel. : +1 303-492-5170 Fax : +1 303-492-5235 email : bruno.giacomazzo@jila.colorado.edu web: http://www.brunogiacomazzo.org
There are only 10 types of people in the world: Those who understand binary, and those who don't
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
Dr. Bruno Giacomazzo JILA - University of Colorado 440 UCB Boulder, CO 80309 USA
Tel. : +1 303-492-5170 Fax : +1 303-492-5235 email : bruno.giacomazzo@jila.colorado.edu web: http://www.brunogiacomazzo.org
---------------------------------------------------------------------- There are only 10 types of people in the world: Those who understand binary, and those who don't ----------------------------------------------------------------------
On 22 Feb 2013, at 02:31, Bruno Giacomazzo bruno.giacomazzo@jila.colorado.edu wrote:
Hi, I did some strong scaling tests on Stampede (https://www.xsede.org/tacc-stampede) with both Whisky and GRHydro (using the development version of ET in both cases). I used a Carpet par file that Roberto DePietri provided me and that he used for similar tests on an Italian cluster (I have attached the GRHydro version, the Whisky one is similar except for using Whisky instead of GRHydro). I used both Intel MPI and Mvapich and I did both pure MPI and MPI/OpenMP runs.
I have attached a text file with my results. The first column is the name of the run (if it starts with mvapich it used mvapich otherwise it used Intel MPI), the second one is the number of cores (option --procs in simfactory), the third one the number of threads (--num-threads), the fourth one the time in seconds spent in CCTK_EVOL, and the fifth one the walltime in seconds (i.e., the total time used by the run as measured on the cluster). I have also attached a couple of figures that show CCTK_EVOL vs #cores and walltime vs #cores (only for Intel MPI runs).
First of all, in pure MPI runs (--num-threads=1) I was unable to run on more than 1024 cores using Intel MPI (the run was just crashing before iteration zero or hanging up). No problem instead when using --num-threads=8 or --num-threads=16. I also noticed that scaling was particularly bad in pure MPI runs and that a lot of time was spent outside CCTK_EVOL (both with Intel MPI and MVAPICH). After speaking with Roberto, I found out that the problem is due to 1D ASCII output (which is active in that parfile) and that makes the runs particularly slow above ~100 cores on this machine. In plot_scaling_walltime_all.pdf I plot also two pure MPI runs, but without 1D ASCII output and the scaling is much better in this case (the time spent in CCTK_EVOL is identical to the case with 1D output and hence I didn't plot them in the other figure). I didn't try using 1D hdf5 output instead, does anyone use it?
According to my tests, --num-threads=16 performs better than --num-threads=8 (which is the current default value in simfactory) and Intel MPI seems to be better than MVAPICH. Is there a particular reason for using 8 instead of 16 threads as the default simfactory value on Stampede?
(replying to an old thread)
I see the same; 16 threads seems to give roughly twice as much performance as 8. Is it possible that there is some issue with cpu/memory socket affinity? There is a tacc_affinity script which the examples say to use, but we don't use it in simfactory. I tried it, and it didn't seem to help. What helped was activating the hwloc thorn; this allowed me to use 8 threads per process, and actually gave slightly better performance than 16, though this might depend on the precise grid structure.
On 2013-09-19, at 14:00 , Ian Hinder ian.hinder@aei.mpg.de wrote:
On 22 Feb 2013, at 02:31, Bruno Giacomazzo bruno.giacomazzo@jila.colorado.edu wrote:
Hi, I did some strong scaling tests on Stampede (https://www.xsede.org/tacc-stampede) with both Whisky and GRHydro (using the development version of ET in both cases). I used a Carpet par file that Roberto DePietri provided me and that he used for similar tests on an Italian cluster (I have attached the GRHydro version, the Whisky one is similar except for using Whisky instead of GRHydro). I used both Intel MPI and Mvapich and I did both pure MPI and MPI/OpenMP runs.
I have attached a text file with my results. The first column is the name of the run (if it starts with mvapich it used mvapich otherwise it used Intel MPI), the second one is the number of cores (option --procs in simfactory), the third one the number of threads (--num-threads), the fourth one the time in seconds spent in CCTK_EVOL, and the fifth one the walltime in seconds (i.e., the total time used by the run as measured on the cluster). I have also attached a couple of figures that show CCTK_EVOL vs #cores and walltime vs #cores (only for Intel MPI runs).
First of all, in pure MPI runs (--num-threads=1) I was unable to run on more than 1024 cores using Intel MPI (the run was just crashing before iteration zero or hanging up). No problem instead when using --num-threads=8 or --num-threads=16. I also noticed that scaling was particularly bad in pure MPI runs and that a lot of time was spent outside CCTK_EVOL (both with Intel MPI and MVAPICH). After speaking with Roberto, I found out that the problem is due to 1D ASCII output (which is active in that parfile) and that makes the runs particularly slow above ~100 cores on this machine. In plot_scaling_walltime_all.pdf I plot also two pure MPI runs, but without 1D ASCII output and the scaling is much better in this case (the time spent in CCTK_EVOL is identical to the case with 1D output and hence I didn't plot them in the other figure). I didn't try using 1D hdf5 output instead, does anyone use it?
According to my tests, --num-threads=16 performs better than --num-threads=8 (which is the current default value in simfactory) and Intel MPI seems to be better than MVAPICH. Is there a particular reason for using 8 instead of 16 threads as the default simfactory value on Stampede?
(replying to an old thread)
I see the same; 16 threads seems to give roughly twice as much performance as 8. Is it possible that there is some issue with cpu/memory socket affinity? There is a tacc_affinity script which the examples say to use, but we don't use it in simfactory. I tried it, and it didn't seem to help. What helped was activating the hwloc thorn; this allowed me to use 8 threads per process, and actually gave slightly better performance than 16, though this might depend on the precise grid structure.
I do all my runs with thorn hwloc, just for this reason. Maybe we should enable it by default.
tacc_affinity should do the same thing, but it may require some options that depend on the number of threads per node. Also, the Intel compiler's startup code may override some of tacc_affinity's settings; maybe one needs to set some environment variables as well. hwloc does its work once the code has started, i.e. it can override all the decisions made by the queuing system, mpirun, tacc_affinity, and the compiler. It also looks at the number of MPI processes and threads running on the current node, so it knows how to do the "right thing".
-erik
users@lists.einsteintoolkit.org