Hi Erik,
> You mention that mvapich has the best performance. Is there any reason
> to use any other MPI implementation?
The version of mvapich required is the latest one. Among the two AMD
clusters I have access to, only Expanse has this version and it is not even
the default one.
The performance with OpenMPI is the exactly the same on one node,
but worse on two nodes.
> Did you check that mvapich is configured correctly? Does it use the
> network efficiently?
How do I do this? Is it on my end, or on the system's end?
> You need to use SystemTopology, or ensure otherwise that the way
> threads and processes are mapped to hardware is reasonable.
I am using SystemTopology.
> What is the ratio of ghost/buffer to actually evolved grid points in your setup?
Is there a quick way to find this out? I am using 14 refinement levels, so
I bet I have a lot of buffer zones. (However, the domain is big.)
> If MPI performance is slow, then the usual way out is to use OpenMP.
> You implied using 4 threads per process; did you try using 8 threads
> per process or more? This will also reduce memory consumption since
> there are fewer ghost zones. Unfortunately, OpenMP multi-threading in
> Carpet is not as efficient as it could be. CarpetX is much better in
> this respect.
I found that on one node, using 4 threads is noticeably faster than
anything else. The admins also recommended this configuration, as it
maps nicely to the hardware. I haven't tried on multiple nodes, though.
> Our way of discretizing equations (high-order methods with 3 ghost
> zones, AMR with buffer zones), combined with having many evolved
> variables, require a lot of storage, and also have a rather high
> parallelization overhead. A few ways out (none are production ready)
> are:
> - Use DGFE instead of finite differences; see e.g. Jonah Miller's PhD
> thesis and the respective McLachlan branch
> - Avoid buffer zones by using an improved time interpolation scheme
> (I've seen papers, I don't know about 3d code)
> - Switch to CarpetX to avoid subcycling in time.
The same parameter file scales "acceptably" on Frontera, so I should be
able to use the same algorithms on Expanse too and obtain some form
of scaling.
> If memory usage is high only on a single node, then this is probably
> caused by a serial code. Known serial pieces of code are ASCII output,
> wave extraction, or the apparent horizon finder. Try disabling these
> to see which (if any) is the culprit.
Carpet prints out a total memory usage that is much less than the
available one, but the simulation still crashes due to out-of-memory error.
Is it possible that one MPI uses much more memory than the others?
> Finally, if OpenMP performance is bad, you can try using only every
> second core and leaving the remainder idle, and see whether this
Yes, this is possible, but I would like to try to find better solutions.
> > - Avoid buffer zones by using an improved time interpolation scheme
> > (I've seen papers, I don't know about 3d code)
> Eloisa, Ian, I and Erik worked on this for a bit quite a while ago. The
> results can be found in branch ianhinder/rkprol:
>
> git clone -b ianhinder/rkprol
https://bitbucket.org/cactuscode/cactusnumerical.git>
> note that this is actually a change to MoL not not so much Carpet. This
> implementation should work, but is not optimized and most likely still
> has (way) too much communication. Based on two presentations at the ET
> meeting in Stockholm:
>
>
https://docs.einsteintoolkit.org/et-docs/ET_Workshop_2015>
> by Bishop Mongwane and Saran Tunyasuvunakool.
If what I report here is a common problem in new AMD systems, it might be
worth seeing if this can find its way into master.
Gabriele