I am currently redesigning the tiling infrastructure, also to allow multithreading via Qthreads instead of OpenMP and to allow for aligning arrays with cache line boundaries. The new approach (different from the current LoopControl) is to choose a fixed tile size, either globally or per loop, and then assign individual tiles to threads. This also works will with DG derivative where the DG element size dictates a granularity for the tile size, and the new efficient tiled derivative operators. Most of this is still in flux. I have seen large efficiency improvements in the RHS calculation, but two puzzling items remain:
(1) It remains more efficient to use MPI than multi-threading for parallelization, at least on regular CPUs. On KNL my results are still somewhat random.
(2) MoL_Add is quite expensive compared to the RHS evaluation.
The main thing that changed since our last round of thorough benchmarks is that CPU became much more powerful while memory bandwidth hasn't. I'm beginning to think that things such as vectorization or parallelization basically don't matter any more if we ensure that we pull data from memory into caches efficiently.
I have not yet collected PAPI statistics.