Bruno
The amount of data output varies greatly between simulations, as well as the intervals between output. And usually, I/O scales very differently than the computation itself. Therefore, it is customary (at least while trying to understand results) to test evolution and output separately. That is, the evolution benchmark would probably output the maximum of rho only, and an I/O benchmark would output (and/or recover) just Minkowski data without time evolution.
Regarding the number of threads: To determine the ideal number of threads to use, we run single-node benchmarks of code that is well parallellised via OpenMP. Code that is not really parallel (e.g. some initial data routines) would distort these results, as would using a large number of cores, since this involves also parallel scalability. Using this "ideal" number of cores, we then run scalability benchmarks to see where parallel scaling breaks down.
For any given physics situation, one then has to strike a compromise:
- non-parallel sections of the code prefer using 1 OpenMP thread
- parallel MPI scaling prefers using as few MPI processes as possible, i.e. using many OpenMP threads
The balance thus shifts with the number of cores -- the more cores you use, the more OpenMP threads you will also want to use to counter-act MPI scaling problems.
As others have said, binding threads to cores and binding memory to processes is also very important, and is visible in particular on single-node benchmarks.
For any benchmark I run, I also look at detailed timer output. The standard Cactus timers are not good enough for this; you will have to use TimerReport or Carpet's timers for this. One interesting quantity is to see what fraction of the time is spent in the actual evolution thorns (not just CCTK_EVOL; this also measures some infrastructure tasks). The other interesting effect to watch is how this time distribution changes as the number of MPI processes increases -- this shows the effect of scaling problems. The latter should e.g. show that ASCII output becomes more time consuming, or that synchronisation or load balancing takes more time.
Finally, I found it very difficult to come up with a "good" benchmark parameter file. Such a benchmark should run both on few and on many cores, should contains all the relevant thorns, should not do I/O, should not contain anything that is known not to scale, should not encounter nans or con2prim problems, should be "close" to actual parameter files that people actually want to use, etc.
I think it's time to create a wiki page for benchmarking! There we could describe (a) the tools available (timers, etc.), (b) the pitfalls to avoid (e.g. measure evolution and I/O separately), and (c) discuss results that we find.
In this case -- I think this means we should improve ASCII output! Writing ASCII files is always slow, but collecting data onto a single process shouldn't be. That's no more than a reduction operation, and with InfiniBand bandwidths of tens of Gigabytes per second, we are very far away from what the hardware performance allows us to do. Let's open a bug report for this. We should either correct this, or should people prominently warn about this.
-erik