I've begun to run the automated tests for the ET on our production machines. Things look very good almost everywhere, except on Stampede, one of the machines that is most important to us. It seems that there are many test failures for GRHydro, and these seem to be caused by segfaults. Does anybody volunteer to investigate?
-erik
On 30 Oct 2014, at 15:03, Erik Schnetter schnetter@cct.lsu.edu wrote:
I've begun to run the automated tests for the ET on our production machines. Things look very good almost everywhere, except on Stampede, one of the machines that is most important to us. It seems that there are many test failures for GRHydro, and these seem to be caused by segfaults. Does anybody volunteer to investigate?
I don't know anything about the problems with GRHydro.
I was having problems a while back with the current simfactory default version of mvapich2, and TACC support suggested I try the mvapich2-x version. The problem I saw was the MPI reductions would hang. I have been using that for a few months with no problems. The required change to the optionlist is:
< MPI_DIR = /opt/apps/intel13/mvapich2/1.9
MPI_DIR = /home1/apps/intel13/mvapich2-x/2.0b MPI_LIB_DIRS = /home1/apps/intel13/mvapich2-x/2.0b/lib64
At least, this was before the changes to the MPI thorn. It's possible that this is no longer enough.
Should we change to this version in simfactory for the release?
Ian
As a side remark, the executables created on Trestles are not runnable. The test cases that nevertheless succeed on Trestles (!) are probably not interesting.
-erik
On Thu, Oct 30, 2014 at 11:34 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 15:03, Erik Schnetter schnetter@cct.lsu.edu wrote:
I've begun to run the automated tests for the ET on our production machines. Things look very good almost everywhere, except on Stampede, one of the machines that is most important to us. It seems that there are many test failures for GRHydro, and these seem to be caused by segfaults. Does anybody volunteer to investigate?
I don't know anything about the problems with GRHydro.
I was having problems a while back with the current simfactory default version of mvapich2, and TACC support suggested I try the mvapich2-x version. The problem I saw was the MPI reductions would hang. I have been using that for a few months with no problems. The required change to the optionlist is:
< MPI_DIR = /opt/apps/intel13/mvapich2/1.9
MPI_DIR = /home1/apps/intel13/mvapich2-x/2.0b MPI_LIB_DIRS = /home1/apps/intel13/mvapich2-x/2.0b/lib64
At least, this was before the changes to the MPI thorn. It's possible that this is no longer enough.
Should we change to this version in simfactory for the release?
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
On 10/30/2014 01:07 PM, Erik Schnetter wrote:
Ian
As a side remark, the executables created on Trestles are not runnable. The test cases that nevertheless succeed on Trestles (!) are probably not interesting.
How does the test succeed if the executable can't run?
Cheers, Steve
-erik
On Thu, Oct 30, 2014 at 11:34 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 15:03, Erik Schnetter schnetter@cct.lsu.edu wrote:
I've begun to run the automated tests for the ET on our production machines. Things look very good almost everywhere, except on Stampede, one of the machines that is most important to us. It seems that there are many test failures for GRHydro, and these seem to be caused by segfaults. Does anybody volunteer to investigate?
I don't know anything about the problems with GRHydro.
I was having problems a while back with the current simfactory default version of mvapich2, and TACC support suggested I try the mvapich2-x version. The problem I saw was the MPI reductions would hang. I have been using that for a few months with no problems. The required change to the optionlist is:
< MPI_DIR = /opt/apps/intel13/mvapich2/1.9
MPI_DIR = /home1/apps/intel13/mvapich2-x/2.0b MPI_LIB_DIRS = /home1/apps/intel13/mvapich2-x/2.0b/lib64
At least, this was before the changes to the MPI thorn. It's possible that this is no longer enough.
Should we change to this version in simfactory for the release?
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
On Thu, Oct 30, 2014 at 2:26 PM, Steven R. Brandt sbrandt@cct.lsu.edu wrote:
On 10/30/2014 01:07 PM, Erik Schnetter wrote:
Ian
As a side remark, the executables created on Trestles are not runnable. The test cases that nevertheless succeed on Trestles (!) are probably not interesting.
How does the test succeed if the executable can't run?
That's the question.
Cactus doesn't test whether the executable runs, it tests whether every generated output file is correct. Maybe there are zero output files.
-erik
Cheers, Steve
-erik
On Thu, Oct 30, 2014 at 11:34 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 15:03, Erik Schnetter schnetter@cct.lsu.edu wrote:
I've begun to run the automated tests for the ET on our production machines. Things look very good almost everywhere, except on Stampede, one of the machines that is most important to us. It seems that there are many test failures for GRHydro, and these seem to be caused by segfaults. Does anybody volunteer to investigate?
I don't know anything about the problems with GRHydro.
I was having problems a while back with the current simfactory default version of mvapich2, and TACC support suggested I try the mvapich2-x version. The problem I saw was the MPI reductions would hang. I have been using that for a few months with no problems. The required change to the optionlist is:
< MPI_DIR = /opt/apps/intel13/mvapich2/1.9
MPI_DIR = /home1/apps/intel13/mvapich2-x/2.0b MPI_LIB_DIRS = /home1/apps/intel13/mvapich2-x/2.0b/lib64
At least, this was before the changes to the MPI thorn. It's possible that this is no longer enough.
Should we change to this version in simfactory for the release?
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 10/30/2014 01:30 PM, Erik Schnetter wrote:
On Thu, Oct 30, 2014 at 2:26 PM, Steven R. Brandt sbrandt@cct.lsu.edu wrote:
On 10/30/2014 01:07 PM, Erik Schnetter wrote:
Ian
As a side remark, the executables created on Trestles are not runnable. The test cases that nevertheless succeed on Trestles (!) are probably not interesting.
How does the test succeed if the executable can't run?
That's the question.
Cactus doesn't test whether the executable runs, it tests whether every generated output file is correct. Maybe there are zero output files.
I thought it also checked exit code.
Cheers, Steve
-erik
Cheers, Steve
-erik
On Thu, Oct 30, 2014 at 11:34 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 15:03, Erik Schnetter schnetter@cct.lsu.edu wrote:
I've begun to run the automated tests for the ET on our production machines. Things look very good almost everywhere, except on Stampede, one of the machines that is most important to us. It seems that there are many test failures for GRHydro, and these seem to be caused by segfaults. Does anybody volunteer to investigate?
I don't know anything about the problems with GRHydro.
I was having problems a while back with the current simfactory default version of mvapich2, and TACC support suggested I try the mvapich2-x version. The problem I saw was the MPI reductions would hang. I have been using that for a few months with no problems. The required change to the optionlist is:
< MPI_DIR = /opt/apps/intel13/mvapich2/1.9
MPI_DIR = /home1/apps/intel13/mvapich2-x/2.0b MPI_LIB_DIRS = /home1/apps/intel13/mvapich2-x/2.0b/lib64
At least, this was before the changes to the MPI thorn. It's possible that this is no longer enough.
Should we change to this version in simfactory for the release?
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 30 Oct 2014, at 19:32, Steven R. Brandt sbrandt@cct.lsu.edu wrote:
On 10/30/2014 01:30 PM, Erik Schnetter wrote:
On Thu, Oct 30, 2014 at 2:26 PM, Steven R. Brandt sbrandt@cct.lsu.edu wrote:
On 10/30/2014 01:07 PM, Erik Schnetter wrote:
Ian
As a side remark, the executables created on Trestles are not runnable. The test cases that nevertheless succeed on Trestles (!) are probably not interesting.
How does the test succeed if the executable can't run?
That's the question.
Cactus doesn't test whether the executable runs, it tests whether every generated output file is correct. Maybe there are zero output files.
I thought it also checked exit code.
mpirun might not be propagating the exit code on that machine.
On 30 Oct 2014, at 20:41, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 19:32, Steven R. Brandt sbrandt@cct.lsu.edu wrote:
On 10/30/2014 01:30 PM, Erik Schnetter wrote:
On Thu, Oct 30, 2014 at 2:26 PM, Steven R. Brandt sbrandt@cct.lsu.edu wrote:
On 10/30/2014 01:07 PM, Erik Schnetter wrote:
Ian
As a side remark, the executables created on Trestles are not runnable. The test cases that nevertheless succeed on Trestles (!) are probably not interesting.
How does the test succeed if the executable can't run?
That's the question.
Cactus doesn't test whether the executable runs, it tests whether every generated output file is correct. Maybe there are zero output files.
I thought it also checked exit code.
mpirun might not be propagating the exit code on that machine.
I think it's worse than that. In WarnLevel.c, in CCTK_VWarn (called by CCTK_Warn), it says
if (level <= error_level) { CCTK_Abort (NULL, 0); }
The second argument to CCTK_Abort is the exit code of the process. So if there is an "error" warning, the process exits with 0 exit code; i.e. success! This happens in several places in this file.
The user guide does not say anything about the exit code of Cactus. I think that if Cactus has a level-0 warning, i.e. an error, then it should exit with a non-zero exit code. Is there a reason to exit "success" in this case?
On 2 Nov 2014, at 17:07, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 20:41, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 19:32, Steven R. Brandt sbrandt@cct.lsu.edu wrote:
On 10/30/2014 01:30 PM, Erik Schnetter wrote:
On Thu, Oct 30, 2014 at 2:26 PM, Steven R. Brandt sbrandt@cct.lsu.edu wrote:
On 10/30/2014 01:07 PM, Erik Schnetter wrote:
Ian
As a side remark, the executables created on Trestles are not runnable. The test cases that nevertheless succeed on Trestles (!) are probably not interesting.
How does the test succeed if the executable can't run?
That's the question.
Cactus doesn't test whether the executable runs, it tests whether every generated output file is correct. Maybe there are zero output files.
I thought it also checked exit code.
mpirun might not be propagating the exit code on that machine.
I think it's worse than that. In WarnLevel.c, in CCTK_VWarn (called by CCTK_Warn), it says
if (level <= error_level) { CCTK_Abort (NULL, 0); }
The second argument to CCTK_Abort is the exit code of the process. So if there is an "error" warning, the process exits with 0 exit code; i.e. success! This happens in several places in this file.
The user guide does not say anything about the exit code of Cactus. I think that if Cactus has a level-0 warning, i.e. an error, then it should exit with a non-zero exit code. Is there a reason to exit "success" in this case?
Further, the test system seems to ignore the fact that Cactus exits with a nonzero exit code. It displays
Cactus exited with error code 1 Please check the logfile...
No files created in test directory
Success: 0 files identical
And in the summary at the end, treats this as a passing test. In this case, there were no test reference files and no files output, because the test does not produce any data, it just aborts if the test fails.
Yes, this should exit with code 1.
-erik
On Sun, Nov 2, 2014 at 11:07 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 20:41, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 19:32, Steven R. Brandt sbrandt@cct.lsu.edu wrote:
On 10/30/2014 01:30 PM, Erik Schnetter wrote:
On Thu, Oct 30, 2014 at 2:26 PM, Steven R. Brandt sbrandt@cct.lsu.edu wrote:
On 10/30/2014 01:07 PM, Erik Schnetter wrote:
Ian
As a side remark, the executables created on Trestles are not runnable. The test cases that nevertheless succeed on Trestles (!) are probably not interesting.
How does the test succeed if the executable can't run?
That's the question.
Cactus doesn't test whether the executable runs, it tests whether every generated output file is correct. Maybe there are zero output files.
I thought it also checked exit code.
mpirun might not be propagating the exit code on that machine.
I think it's worse than that. In WarnLevel.c, in CCTK_VWarn (called by CCTK_Warn), it says
if (level <= error_level) { CCTK_Abort (NULL, 0); }
The second argument to CCTK_Abort is the exit code of the process. So if there is an "error" warning, the process exits with 0 exit code; i.e. success! This happens in several places in this file.
The user guide does not say anything about the exit code of Cactus. I think that if Cactus has a level-0 warning, i.e. an error, then it should exit with a non-zero exit code. Is there a reason to exit "success" in this case?
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Ian
This new MPI version leads to problems running the benchmarks, and runs at half the speed. (This test was on a single node.)
-erik
On Thu, Oct 30, 2014 at 11:34 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 15:03, Erik Schnetter schnetter@cct.lsu.edu wrote:
I've begun to run the automated tests for the ET on our production machines. Things look very good almost everywhere, except on Stampede, one of the machines that is most important to us. It seems that there are many test failures for GRHydro, and these seem to be caused by segfaults. Does anybody volunteer to investigate?
I don't know anything about the problems with GRHydro.
I was having problems a while back with the current simfactory default version of mvapich2, and TACC support suggested I try the mvapich2-x version. The problem I saw was the MPI reductions would hang. I have been using that for a few months with no problems. The required change to the optionlist is:
< MPI_DIR = /opt/apps/intel13/mvapich2/1.9
MPI_DIR = /home1/apps/intel13/mvapich2-x/2.0b MPI_LIB_DIRS = /home1/apps/intel13/mvapich2-x/2.0b/lib64
At least, this was before the changes to the MPI thorn. It's possible that this is no longer enough.
Should we change to this version in simfactory for the release?
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
On 30 Oct 2014, at 21:55, Erik Schnetter schnetter@cct.lsu.edu wrote:
Ian
This new MPI version leads to problems running the benchmarks, and runs at half the speed. (This test was on a single node.)
Ouch. I didn't notice that with my runs; either I wasn't paying attention or it didn't happen there. I will check the next time I run on stampede.
-erik
On Thu, Oct 30, 2014 at 11:34 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 15:03, Erik Schnetter schnetter@cct.lsu.edu wrote:
I've begun to run the automated tests for the ET on our production machines. Things look very good almost everywhere, except on Stampede, one of the machines that is most important to us. It seems that there are many test failures for GRHydro, and these seem to be caused by segfaults. Does anybody volunteer to investigate?
I don't know anything about the problems with GRHydro.
I was having problems a while back with the current simfactory default version of mvapich2, and TACC support suggested I try the mvapich2-x version. The problem I saw was the MPI reductions would hang. I have been using that for a few months with no problems. The required change to the optionlist is:
< MPI_DIR = /opt/apps/intel13/mvapich2/1.9
MPI_DIR = /home1/apps/intel13/mvapich2-x/2.0b MPI_LIB_DIRS = /home1/apps/intel13/mvapich2-x/2.0b/lib64
At least, this was before the changes to the MPI thorn. It's possible that this is no longer enough.
Should we change to this version in simfactory for the release?
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
On 31 Oct 2014, at 09:52, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 21:55, Erik Schnetter schnetter@cct.lsu.edu wrote:
Ian
This new MPI version leads to problems running the benchmarks, and runs at half the speed. (This test was on a single node.)
Ouch. I didn't notice that with my runs; either I wasn't paying attention or it didn't happen there. I will check the next time I run on stampede.
In my production simulations, I see a speed drop of 20% going from the current simfactory and stampede default of mvapich2/1.9 to mvapich2-x/2.0b as suggested by the TACC admins. However, since the original simulations were hanging, I'm not sure which is better!
Looking at the timers, prolongate is taking 868.4s with 1.9 and 1313.6s with 2.0b. Sync is about the same speed on both, as are computational functions. The -x suffix on the mvapich version seems to indicate that it supports the MICs; maybe there is some trade-off that is made.
-erik
On Thu, Oct 30, 2014 at 11:34 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 15:03, Erik Schnetter schnetter@cct.lsu.edu wrote:
I've begun to run the automated tests for the ET on our production machines. Things look very good almost everywhere, except on Stampede, one of the machines that is most important to us. It seems that there are many test failures for GRHydro, and these seem to be caused by segfaults. Does anybody volunteer to investigate?
I don't know anything about the problems with GRHydro.
I was having problems a while back with the current simfactory default version of mvapich2, and TACC support suggested I try the mvapich2-x version. The problem I saw was the MPI reductions would hang. I have been using that for a few months with no problems. The required change to the optionlist is:
< MPI_DIR = /opt/apps/intel13/mvapich2/1.9
MPI_DIR = /home1/apps/intel13/mvapich2-x/2.0b MPI_LIB_DIRS = /home1/apps/intel13/mvapich2-x/2.0b/lib64
At least, this was before the changes to the MPI thorn. It's possible that this is no longer enough.
Should we change to this version in simfactory for the release?
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
I don't think there's a trade-off involved. I attended one or two presentations by mvapich developers, and the additional complexity from handling MICs comes from correctly (and efficiently) routing data between CPUs, MICs, and network interfaces within a node, where multiple paths may exist, and where these paths have different performance for different message sizes.
The slow-down may be caused by different default parameter settings in 1.9 and 2.0; maybe certain environment variables could change performance.
In any way, my earlier tests of 2.0 on Stampede were probably tainted by problems with PETSc and HDF5, and we should repeat them.
-erik
On Wed, Nov 5, 2014 at 5:04 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 31 Oct 2014, at 09:52, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 21:55, Erik Schnetter schnetter@cct.lsu.edu wrote:
Ian
This new MPI version leads to problems running the benchmarks, and runs at half the speed. (This test was on a single node.)
Ouch. I didn't notice that with my runs; either I wasn't paying attention or it didn't happen there. I will check the next time I run on stampede.
In my production simulations, I see a speed drop of 20% going from the current simfactory and stampede default of mvapich2/1.9 to mvapich2-x/2.0b as suggested by the TACC admins. However, since the original simulations were hanging, I'm not sure which is better!
Looking at the timers, prolongate is taking 868.4s with 1.9 and 1313.6s with 2.0b. Sync is about the same speed on both, as are computational functions. The -x suffix on the mvapich version seems to indicate that it supports the MICs; maybe there is some trade-off that is made.
-erik
On Thu, Oct 30, 2014 at 11:34 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 30 Oct 2014, at 15:03, Erik Schnetter schnetter@cct.lsu.edu wrote:
I've begun to run the automated tests for the ET on our production machines. Things look very good almost everywhere, except on Stampede, one of the machines that is most important to us. It seems that there are many test failures for GRHydro, and these seem to be caused by segfaults. Does anybody volunteer to investigate?
I don't know anything about the problems with GRHydro.
I was having problems a while back with the current simfactory default version of mvapich2, and TACC support suggested I try the mvapich2-x version. The problem I saw was the MPI reductions would hang. I have been using that for a few months with no problems. The required change to the optionlist is:
< MPI_DIR = /opt/apps/intel13/mvapich2/1.9
MPI_DIR = /home1/apps/intel13/mvapich2-x/2.0b MPI_LIB_DIRS = /home1/apps/intel13/mvapich2-x/2.0b/lib64
At least, this was before the changes to the MPI thorn. It's possible that this is no longer enough.
Should we change to this version in simfactory for the release?
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
users@lists.einsteintoolkit.org