* intel19 + intel mpi : great 1 node speed, poor scaling to >1 node, poor weak scaling * intel19 + openmpi 4 : same as above * gcc 10 + openmpi 4 : 40% slower 1 node speed, poor scaling to >1 node, poor weak scaling * gcc 9 + openmpi 3.1.6 : 40% slower 1 node speed, good scaling to >1 node, acceptable weak scaling * using intel19 + openmpi 4 executable but with gcc9/openmpi 3.1.6 modules loaded at runtime : great 1 node speed, good scaling to >1 node, acceptable weak scaling
"I think I can install openmpi/3.1.6 with intel compilers. I have to go back and check but I think the main difference is we are using ibverbs on the openmpi/3.1.6 build and ucx on the openmpi/4.0.4. For most codes ucx has been the faster option but in your case it seems different. I will let you know once the compilers are in place."
Hello,
Last week I opened a PR to add the configuration filesfor Expanse to simfactory. Expanse is an example ofthe new generation of AMD supercomputers. Others areAnvil, one of the other new XSEDE machines, or Puma,the newest cluster at The University of Arizona.
I have some experience with Puma and Expanse andI would like to share some thoughts, some of which comefrom interacting with the admins of Expanse. The problemis that I am finding terrible multi-node performance on boththese machines, and I don't know if this will be a commonthread among new AMD clusters.
These supercomputers have similar characteristics.
First, they have very high cores/node count (typically128/node) but low memory per core (typically 2 GB / core).In these conditions, it is very easy to have a job killed bythe OOM daemon. My suspicion is that it is rank 0 thatgoes out of memory, and the entire run is aborted.
Second, depending on the MPI implementation, MPI collectiveoperations can be extremely expensive. I was told thatthe best implementation is mvapich 2.3.6 (at the moment).This seems to be due to the high core count.
I found that the code does not scale well. This is possiblyrelated to the previous point. If your job can fit on a single node,it will run wonderfully. However, when you perform the samesimulation on two nodes, the code will actually be slower.This indicates that there's no strong scaling at all from1 node to 2 (128 to 256 cores, or 32 to 64 MPI ranks).Using mvapich 2.3.6 improves the situation, but it is stillfaster to use fewer nodes.
(My benchmark is a par file I've tested extensively on Frontera)
I am working with Expanse's support staff to see what we cando, but I wonder if anyone has had a positive experience withthis architecture and has some tips to share.
Gabriele
_______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users