Helvi, Miguel
I find that many (most)? users of the Einstein Toolkit define their problem size based on physics and accuracy requirements, and then adapt the number of nodes they uses to make this problem run most efficiently. This usually requires a certain (fixed) amount of work per core. Correspondingly, an informative benchmarking chart is a "weak scaling test". Strong scalability is also interesting, but is more difficult to interpret in the presence of adaptive mesh refinement with a complex system of equations since using too few nodes leads to out-of-memory situations, and using too many nodes quickly leads to inefficiencies if you use higher-order derivatives.
Efficiency depends also on the number of OpenMP threads vs. the number of MPI processes, and (obviously) the amount of I/O you're doing, which is governed by very different characteristics (the file system used, number of file servers, etc.) We thus typically benchmark weak scalability on setups that perform very little I/O, and where we exclude initial data generation.
When setting up a benchmark, it is important to monitor as many performance characteristics as possible to ensure that one isn't limited by I/O, or by any other "accidental" feature that could easily be removed in a production simulation such as horizon finding, constraint evaluation, too frequent regridding, etc.
It is surprisingly easy to accidentally enable a feature in the Einstein Toolkit that adversely affects performance, in particular when running on 1000+ cores. Correspondingly, one has to take great care when defining a benchmark that no such features are present. Timer output will be valuable here to understand how much
A QC-0 setup should be a good test case if you want to study BBH scenarios. Or maybe GW150914 <
http://einsteintoolkit.org/gallery/bbh/index.html> would be more interesting and relevant? You would obviously crank up the resolution to turn this into a weak scaling test. In my experience, running for a few minutes (less than ten) after initial data setup will suffice to get good numbers.
The benchmarks I am running are usually simpler since I tend to focus on a particular features that I want to optimize (e.g. RHS evaluation, grid structure management). I think it's been some time since we ran production-scale BBH benchmarks (with puncture tracking, regridding tracking the black holes, etc.), I believe we focussed on hydrodynamics benchmarks recently.
The final quantity against which I report performance is usually "grid point updates per second", plotted against the number of nodes used. This is the time required to evaluate the RHS for a single grid point, amortized over and including all overhead such as scheduling, regridding, prolongation, synchronization, etc. (Since this measures time, smaller numbers are better.) On a fast machine, and if the overhead is low, this number can be as low as a few microseconds. Mesh refinement and parallelization will add a non-negligible overhead to this, and a weak scaling test will show when the parallelization overhead becomes prohibitive.
-erik