Hello,
Two days ago, I opened a PR to the simfactory repo to add Expanse, the newest machine at the San Diego Supercomputing Center, based on AMD Epyc "Rome" CPUs and part of XSEDE. In the meantime, I realized that some tests are failing miserably, but I couldn't figure out why.
Before I describe what I found, let me start with a side node on AMD compilers.
<side node>
There are four compilers available on Expanse: GNU, Intel, AMD, and PGI. I did not touch the PGI compilers. I briefly tried (and failed) to compile with the AMD compilers (aocc and flang). I did not try hard, and it seems that most of the libraries on Expanse are compiled with gcc anyways.
A first step to support these compilers is adding the lines:
elif test "`$F90 --version 2>&1 | grep AMD`" ; then LINUX_F90_COMP=AMD else
elif test "`$CC --version 2>&1 | grep AMD`" ; then LINUX_C_COMP=AMD fi
elif test "`$CC --version 2>&1 | grep AMD`" ; then LINUX_CXX_COMP=AMD fi
in the obvious places in flesh/lib/make/known-architecture/linux.
</side node>
I successfully compiled the Einstein Toolkit with - gcc 10.2.0 and OpenMPI 4.0.4 - gcc 9.2.0 and OpenMPI 4.0.4 - intel 2019 and Intel MPI 2019
I noticed that some tests, like ADMMass/tov_carpet.par, gave completely incorrect results. For example, the expected value is 1.3, but I would find 1.6.
I disabled all the optimizations, but the test would keep failing. At the end, I noticed that if I ran with 8/16/32 MPI processes per node, and the corresponding number of OpenMP threads (128/N_MPI), the test would fail, but if I ran with 4/2/1 MPI processes, the test would pass.
Most of my experiments were with gcc 10, but the test fails also with the Intel suite.
I tried increasing the OMP_STACK_SIZE to a very large value, but it didn't help.
Any idea of what the problem might be?
Gabriele
Hello Gabriele,
Thank you for contributing these.
The test suites are quick running parfiles with small grids, so running them on large numbers of MPI ranks (they are designed for 1 or 2 MPI ranks) can lead to unexpected situations (such as an MPI rank having no grid points at all).
Generally, if the tests work for 1,2,4 ranks (4 being the largest number of procs requested by any test.ccl file) then this is sufficient.
In principle even running on more MPI ranks should work, so if you know which tests fail with the larger number of MPI ranks and were to list them in a ticket, maybe someone could look into this.
Note that you can undersubscribe compute node, in particular for tests, if you do not need / want to use all cores.
Can you create a pull request for the "linux" architecture file with the changes for the AMD compiler you found, please? So far it sees you mostly only changed the detection part, does it then not also require some changes in the "set values" part of the file? Eg default values for optimization, preprocessor or so?
Yours, Roland
Hello,
Two days ago, I opened a PR to the simfactory repo to add Expanse, the newest machine at the San Diego Supercomputing Center, based on AMD Epyc "Rome" CPUs and part of XSEDE. In the meantime, I realized that some tests are failing miserably, but I couldn't figure out why.
Before I describe what I found, let me start with a side node on AMD compilers.
<side node>
There are four compilers available on Expanse: GNU, Intel, AMD, and PGI. I did not touch the PGI compilers. I briefly tried (and failed) to compile with the AMD compilers (aocc and flang). I did not try hard, and it seems that most of the libraries on Expanse are compiled with gcc anyways.
A first step to support these compilers is adding the lines:
elif test "`$F90 --version 2>&1 | grep AMD`" ; then LINUX_F90_COMP=AMD else
elif test "`$CC --version 2>&1 | grep AMD`" ; then LINUX_C_COMP=AMD fi
elif test "`$CC --version 2>&1 | grep AMD`" ; then LINUX_CXX_COMP=AMD fi
in the obvious places in flesh/lib/make/known-architecture/linux.
</side node>
I successfully compiled the Einstein Toolkit with
- gcc 10.2.0 and OpenMPI 4.0.4
- gcc 9.2.0 and OpenMPI 4.0.4
- intel 2019 and Intel MPI 2019
I noticed that some tests, like ADMMass/tov_carpet.par, gave completely incorrect results. For example, the expected value is 1.3, but I would find 1.6.
I disabled all the optimizations, but the test would keep failing. At the end, I noticed that if I ran with 8/16/32 MPI processes per node, and the corresponding number of OpenMP threads (128/N_MPI), the test would fail, but if I ran with 4/2/1 MPI processes, the test would pass.
Most of my experiments were with gcc 10, but the test fails also with the Intel suite.
I tried increasing the OMP_STACK_SIZE to a very large value, but it didn't help.
Any idea of what the problem might be?
Gabriele
Hi Roland,
thanks for your answer.
The test suites are quick running parfiles with small grids, so running
them on large numbers of MPI ranks (they are designed for 1 or 2 MPI ranks) can lead to unexpected situations (such as an MPI rank having no grid points at all).
Generally, if the tests work for 1,2,4 ranks (4 being the largest
number of procs requested by any test.ccl file) then this is sufficient.
Frontera and Stampede2 use 24/28 MPI processes, but the tests still pass. I am particularly looking at the test ADMMass/tov_carpet.par, where the numbers are off, but no error is thrown. Another example is Exact/de_Sitter.par. Other tests do fail because of Carpet errors, which might be what you are describing.
Can you create a pull request for the "linux" architecture file with
the changes for the AMD compiler you found, please? So far it sees you mostly only changed the detection part, does it then not also require some changes in the "set values" part of the file? Eg default values for optimization, preprocessor or so?
Where is the repo?
I am not too familiar with what that file is supposed to set. But, I only changed what was needed to at least start the compilation.
Gabriele
On Wed, Aug 18, 2021 at 8:20 AM Roland Haas rhaas@illinois.edu wrote:
Hello Gabriele,
Thank you for contributing these.
The test suites are quick running parfiles with small grids, so running them on large numbers of MPI ranks (they are designed for 1 or 2 MPI ranks) can lead to unexpected situations (such as an MPI rank having no grid points at all).
Generally, if the tests work for 1,2,4 ranks (4 being the largest number of procs requested by any test.ccl file) then this is sufficient.
In principle even running on more MPI ranks should work, so if you know which tests fail with the larger number of MPI ranks and were to list them in a ticket, maybe someone could look into this.
Note that you can undersubscribe compute node, in particular for tests, if you do not need / want to use all cores.
Can you create a pull request for the "linux" architecture file with the changes for the AMD compiler you found, please? So far it sees you mostly only changed the detection part, does it then not also require some changes in the "set values" part of the file? Eg default values for optimization, preprocessor or so?
Yours, Roland
Hello,
Two days ago, I opened a PR to the simfactory repo to add Expanse, the newest machine at the San Diego Supercomputing Center, based on AMD Epyc "Rome" CPUs and part of XSEDE. In the meantime, I realized that some tests are failing miserably, but I couldn't figure out why.
Before I describe what I found, let me start with a side node on AMD compilers.
<side node>
There are four compilers available on Expanse: GNU, Intel, AMD, and PGI. I did not touch the PGI compilers. I briefly tried (and failed) to
compile
with the AMD compilers (aocc and flang). I did not try hard, and it seems that most of the libraries on Expanse are compiled with gcc anyways.
A first step to support these compilers is adding the lines:
elif test "`$F90 --version 2>&1 | grep AMD`" ; then LINUX_F90_COMP=AMD else
elif test "`$CC --version 2>&1 | grep AMD`" ; then LINUX_C_COMP=AMD fi
elif test "`$CC --version 2>&1 | grep AMD`" ; then LINUX_CXX_COMP=AMD fi
in the obvious places in flesh/lib/make/known-architecture/linux.
</side node>
I successfully compiled the Einstein Toolkit with
- gcc 10.2.0 and OpenMPI 4.0.4
- gcc 9.2.0 and OpenMPI 4.0.4
- intel 2019 and Intel MPI 2019
I noticed that some tests, like ADMMass/tov_carpet.par, gave completely incorrect results. For example, the expected value is 1.3, but I would find 1.6.
I disabled all the optimizations, but the test would keep failing. At the end, I noticed that if I ran with 8/16/32 MPI processes per node, and the corresponding number of OpenMP threads (128/N_MPI), the test would fail, but if I ran with 4/2/1 MPI processes, the test would pass.
Most of my experiments were with gcc 10, but the test fails also with the Intel suite.
I tried increasing the OMP_STACK_SIZE to a very large value, but it didn't help.
Any idea of what the problem might be?
Gabriele
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://pgp.mit.edu .
Hello Gabriele,
Frontera and Stampede2 use 24/28 MPI processes, but the tests still pass. I am particularly looking at the test ADMMass/tov_carpet.par, where the numbers are off, but no error is thrown. Another example is Exact/de_Sitter.par.
Hmm. Would be interesting to see if the same error happens eg on a workstation where one compiles with gcc but then runs with say 32 MPI ranks (even if that oversubscribes the workstation). It is possible. of course, that there is a long surviving race condition or bug in these thorns.
We also have some issues with some tests failing on Blue Waters but I have no idea why and it is not reproducible on any other system and is deep in some F77 code.
Other tests do fail because of Carpet errors, which might be what you are describing.
ok.
Can you create a pull request for the "linux" architecture file with the changes for the AMD compiler you found, please? So far it sees you mostly only changed the detection part, does it then not also require some changes in the "set values" part of the file? Eg default values for optimization, preprocessor or so?
Where is the repo?
The repository is the Cactus "flesh" repo. You can find out which one is is using git. eg:
cd lib/make/known-architectures git remote -v
which in this case reports:
https://bitbucket.org/cactuscode/cactus.git
The directory of the checkout can be obtained by either looking at the symbolic links, or via:
cd lib/make/known-architectures pwd -P
which shows the full, all links resolved path to the current working directory.
I am not too familiar with what that file is supposed to set. But, I only changed what was needed to at least start the compilation.
If you are using an option list then I would have hoped that nothing in the file requires changing. You might expect to get a warning about this being an untested architecture, but that should have been it. So in your case it actually failed b/c it could not identify the compiler? That is somewhat annoying.... and indeed the case looking at the fragment:
if ! test "x$LINUX_C_COMP" = "xunknown" ; then echo "Internal error: did not expect Linux C compiler to be $LINUX_C_COMP" exit 2 fi
What would need adjusting would be the cases statements and you should add options similar to eg what is being provided for GNU. Something like:
: ${CFLAGS='-std=gnu99'} : ${C_OPTIMISE_FLAGS='-O3'} CC_VERSION="`$CC -v 2>&1 | grep -i "AOCC version" | head -n1`" : ${C_OPENMP_FLAGS='-fopenmp'}
where I am not adding any support for an AOCC compiler that does not even support OpenMP. I do not think we will find any such compiler nowadays (since we require compilers to support eg c++11 I find it very unlikely that we would find a compiler suite that does support C++11 but not OpenMP).
The colon ":" is the POSIX compliant name for "true" and we do not really care about it, only about the ${FOO=bar} default value variable assignment that the shell performs before executing "true".
Yours, Roland
Hi Roland,
Hmm. Would be interesting to see if the same error happens eg on a
workstation where one compiles with gcc but then runs with say 32 MPI ranks (even if that oversubscribes the workstation). It is possible. of course, that there is a long surviving race condition or bug in these thorns.
I realized I was not comparing the same things. In fact, on Frontera I ran the tests with up to 2 MPI processes. When I restrict to 1/2 MPI processes, almost all tests pass on Expanse, so I guess that mine was a false alarm and everything is all right. I can upload the test results on the repo.
So in your case it actually failed b/c it could not identify the compiler?
Yes, correct.
I set all the other variables in the option list, but at the end I didn't end up compiling with aocc because of some issues with external libraries (if I remember correctly).
I can add the code for detecting aocc, but I would leave everything else to someone that knows exactly what variables should be defined and how.
Gabriele
On Wed, Aug 18, 2021 at 11:00 AM Roland Haas rhaas@illinois.edu wrote:
Hello Gabriele,
Frontera and Stampede2 use 24/28 MPI processes, but the tests still pass. I am particularly looking at the test ADMMass/tov_carpet.par, where the numbers are off, but no error is thrown. Another example is Exact/de_Sitter.par.
Hmm. Would be interesting to see if the same error happens eg on a workstation where one compiles with gcc but then runs with say 32 MPI ranks (even if that oversubscribes the workstation). It is possible. of course, that there is a long surviving race condition or bug in these thorns.
We also have some issues with some tests failing on Blue Waters but I have no idea why and it is not reproducible on any other system and is deep in some F77 code.
Other tests do fail because of Carpet errors, which might be what you are describing.
ok.
Can you create a pull request for the "linux" architecture file with the changes for the AMD compiler you found, please? So far it sees you mostly only changed the detection part, does it then not also require some changes in the "set values" part of the file? Eg default values for optimization, preprocessor or so?
Where is the repo?
The repository is the Cactus "flesh" repo. You can find out which one is is using git. eg:
cd lib/make/known-architectures git remote -v
which in this case reports:
https://bitbucket.org/cactuscode/cactus.git
The directory of the checkout can be obtained by either looking at the symbolic links, or via:
cd lib/make/known-architectures pwd -P
which shows the full, all links resolved path to the current working directory.
I am not too familiar with what that file is supposed to set. But, I only changed what was needed to at least start the compilation.
If you are using an option list then I would have hoped that nothing in the file requires changing. You might expect to get a warning about this being an untested architecture, but that should have been it. So in your case it actually failed b/c it could not identify the compiler? That is somewhat annoying.... and indeed the case looking at the fragment:
if ! test "x$LINUX_C_COMP" = "xunknown" ; then echo "Internal error: did not expect Linux C compiler to be $LINUX_C_COMP" exit 2 fi
What would need adjusting would be the cases statements and you should add options similar to eg what is being provided for GNU. Something like:
: ${CFLAGS='-std=gnu99'} : ${C_OPTIMISE_FLAGS='-O3'} CC_VERSION="`$CC -v 2>&1 | grep -i "AOCC version" | head -n1`" : ${C_OPENMP_FLAGS='-fopenmp'}
where I am not adding any support for an AOCC compiler that does not even support OpenMP. I do not think we will find any such compiler nowadays (since we require compilers to support eg c++11 I find it very unlikely that we would find a compiler suite that does support C++11 but not OpenMP).
The colon ":" is the POSIX compliant name for "true" and we do not really care about it, only about the ${FOO=bar} default value variable assignment that the shell performs before executing "true".
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://pgp.mit.edu .
Hello Miguel,
I realized I was not comparing the same things. In fact, on Frontera I ran the tests with up to 2 MPI processes. When I restrict to 1/2 MPI processes, almost all tests pass on Expanse, so I guess that mine was a false alarm and everything is all right. I can upload the test results on the repo.
Oha. I seem to have misunderstood your question before. The repository for the test results shown on
https://einsteintoolkit.org/testsuite_results/index.php
is:
https://bitbucket.org/einsteintoolkit/testsuite_results
(as shown at the top of that page).
However that is only for the release ET version.
There is no repository for testsuite results of the development (trunk) version.
I can add the code for detecting aocc, but I would leave everything else to someone that knows exactly what variables should be defined and how.
Having a pull request with what you have would greatly simplify anyone else continuing from there since they would have a working staring point.
Did you already create a pull request and ticket for the simfactory files needed to use Expanse?
Right now I see a pull request on
https://bitbucket.org/simfactory/simfactory2/pull-requests/
but no ticket on
https://bitbucket.org/einsteintoolkit/tickets/issues/
yet.
Yours, Roland
Hi Roland,
you probably understood the problem correctly (test fail with 32 MPI processes). But, the reason I asked the question in the first place was wrong, since I thought that the same test was passing on Frontera and Stampede, but I was actually running a different test.
I uploaded the results of the tests to the restuite_results repo, and I opened a PR and a ticket for adding Expanse. I will create a PR to add basic support to aocc too.
Gabriele
On Wed, Aug 18, 2021 at 12:13 PM Roland Haas rhaas@illinois.edu wrote:
Hello Miguel,
I realized I was not comparing the same things. In fact, on Frontera I
ran
the tests with up to 2 MPI processes. When I restrict to 1/2 MPI processes, almost all tests pass on Expanse, so I guess that mine was a false alarm and everything is all right. I can upload the test results on the repo.
Oha. I seem to have misunderstood your question before. The repository for the test results shown on
https://einsteintoolkit.org/testsuite_results/index.php
is:
https://bitbucket.org/einsteintoolkit/testsuite_results
(as shown at the top of that page).
However that is only for the release ET version.
There is no repository for testsuite results of the development (trunk) version.
I can add the code for detecting aocc, but I would leave everything else
to
someone that knows exactly what variables should be defined and how.
Having a pull request with what you have would greatly simplify anyone else continuing from there since they would have a working staring point.
Did you already create a pull request and ticket for the simfactory files needed to use Expanse?
Right now I see a pull request on
https://bitbucket.org/simfactory/simfactory2/pull-requests/
but no ticket on
https://bitbucket.org/einsteintoolkit/tickets/issues/
yet.
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://pgp.mit.edu .
users@lists.einsteintoolkit.org