Hi, I have sent the pull request with the optionlist for Stampede - KNL on Bitbucket simfactory repo. I have tested this with a couple of thornlists including the einsteintoolkit.th and GW150914.th. This is still in experimental stage and so would be great if someone could also test it.
Working on benchmarking the performance on Stampede KNL, I was able to do some test runs using the GW150914 simulation. However, I have been running into some issues with it.
1. I tried running QC0 simulation on both Stampede SandyBridge and KNL. While it runs fine on Stampede but it crashes on KNL with this error -
while executing schedule bin BoundaryConditions, routine RotatingSymmetry180::Rot180_ApplyBC in thorn RotatingSymmetry180, file /work/04082/tg833814/Cactus_ETK_dev/arrangements/CactusNumerical/RotatingSymmetry180/src/rotatingsymmetry180.c:460: -> TAT/Slab can only be used if there is a single local component per MPI process TACC: MPI job exited with code: 134 I looked up at previous tickets and found the solution to increase the number of cores. But if the same simulation can be run on stampede on 64 cores, why does it require higher number of cores on KNL? Or is it some other issue?
2. I was able to run GW150914 on development queue (68 cores) and the speeds on Stampede were around 12.9M while that on KNL goes around 2.4M. To understand the reason for such small speeds, I tried running this on higher number of cores on Stampede (128) and it runs at speed of around 20.9M (tested the run for 12 hours). However, on doing the same in normal queue in KNL, the simulation crashes after a couple of iterations on KNL with some segmentation fault error. Also, before crashing, the speed on KNL is around 4.2M. I have attached the error file of the simulation.
Could someone please look at this? Let me know if you need any other information.
Thanks
.............................
Bhavesh Khamesra
Graduate Student
Centre of Relativistic Astrophysics
Georgia Institute of Technology
Bhavesh:
- I was able to run GW150914 on development queue (68 cores) and the speeds on Stampede were around 12.9M while that on KNL goes around 2.4M. To understand the reason for such small speeds, I tried running this on higher number of cores on Stampede (128) and it runs at speed of around 20.9M (tested the run for 12 hours). However, on doing the same in normal queue in KNL, the simulation crashes after a couple of iterations on KNL with some segmentation fault error. Also, before crashing, the speed on KNL is around 4.2M. I have attached the error file of the simulation.
FYI, I have seen the same with my code: it runs (very slowly) on a single KNL node, but it crashes immediately when running on more than 1 node. Curiously, I do not have the same problem on Marconi KNL, an Italian cluster with the same architecture/interconnect as Stampede KNL.
Best wishes,
David
Hi Dr Radice, Thanks for the reply. What is the general speed you get on Marconi KNL for GW150914 case (or for any other simulation)?
Regards
.............................
Bhavesh Khamesra
Graduate Student
Centre of Relativistic Astrophysics
Georgia Institute of Technology
________________________________ From: David Radice david.e.pi.3.14@gmail.com Sent: Wednesday, May 3, 2017 2:54:51 PM To: Khamesra, Bhavesh Cc: users@einsteintoolkit.org Subject: Re: [Users] Benchmarking
Bhavesh:
- I was able to run GW150914 on development queue (68 cores) and the speeds on Stampede were around 12.9M while that on KNL goes around 2.4M. To understand the reason for such small speeds, I tried running this on higher number of cores on Stampede (128) and it runs at speed of around 20.9M (tested the run for 12 hours). However, on doing the same in normal queue in KNL, the simulation crashes after a couple of iterations on KNL with some segmentation fault error. Also, before crashing, the speed on KNL is around 4.2M. I have attached the error file of the simulation.
FYI, I have seen the same with my code: it runs (very slowly) on a single KNL node, but it crashes immediately when running on more than 1 node. Curiously, I do not have the same problem on Marconi KNL, an Italian cluster with the same architecture/interconnect as Stampede KNL.
Best wishes,
David
Hi Dr Radice, Thanks for the reply. What is the general speed you get on Marconi KNL for GW150914 case (or for any other simulation)?
I have not tried a BBH run myself, but Eloisa reported, on this mailing list, that
I now obtain a runspeed on a Marconi KNL node which is around 80% of a Xeon E5 v4.
You can find more details in the mailing list archives.
Cheers,
David
Bhavesh
To be exact, the remedy for this particular Slab error is not to use more cores, but to use more MPI processes. You can keep the number of cores constant if you reduce the number of OpenMP threads per MPI process.
Given that you are benchmarking, you should anyway experiment with these parameters, as performance can crucially depend on them. Usually, using fewer threads and more processes is more efficient for small core counts.
Finally, only comparing the overall run time is not sufficient to make a statement about performance. Each run has several "tuning knobs", and choosing the right values for these is important to achieve good performance. Using the default settings will often lead to quite poor performance. Cactus timer output as well as experience with performing runs on HPC systems is indispensable to get good performance.
-erik
On Tue, May 2, 2017 at 5:09 PM, Khamesra, Bhavesh < bhaveshkhamesra@gatech.edu> wrote:
Hi, I have sent the pull request with the optionlist for Stampede - KNL on Bitbucket simfactory repo. I have tested this with a couple of thornlists including the einsteintoolkit.th and GW150914.th. This is still in experimental stage and so would be great if someone could also test it.
Working on benchmarking the performance on Stampede KNL, I was able to do some test runs using the GW150914 simulation. However, I have been running into some issues with it.
- I tried running QC0 simulation on both Stampede SandyBridge and KNL.
While it runs fine on Stampede but it crashes on KNL with this error -
while executing schedule bin BoundaryConditions, routine Rota tingSymmetry180::Rot180_ApplyBC in thorn RotatingSymmetry180, file /work/04082/tg833814/Cactus_ETK_dev/arrangements/CactusNumerical/ RotatingSymmetry180/src/rotatingsymmetry180.c:460:
-> TAT/Slab can only be used if there is a single local component per MPI process TACC: MPI job exited with code: 134 I looked up at previous tickets and found the solution to increase the number of cores. But if the same simulation can be run on stampede on 64 cores, why does it require higher number of cores on KNL? Or is it some other issue?
- I was able to run GW150914 on development queue (68 cores) and the
speeds on Stampede were around 12.9M while that on KNL goes around 2.4M. To understand the reason for such small speeds, I tried running this on higher number of cores on Stampede (128) and it runs at speed of around 20.9M (tested the run for 12 hours). However, on doing the same in normal queue in KNL, the simulation crashes after a couple of iterations on KNL with some segmentation fault error. Also, before crashing, the speed on KNL is around 4.2M. I have attached the error file of the simulation.
Could someone please look at this? Let me know if you need any other information.
Thanks
.............................
Bhavesh Khamesra
Graduate Student
Centre of Relativistic Astrophysics
Georgia Institute of Technology
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Hi Erik, Thanks for the reply. I tried playing with num-threads options in machine files and was able to run the QC0 on development node. Reducing the num-threads to 17 keeping number of cores to 64 showed some increase in the speed but it is still quite low - around 13-16M/hour compared the 55-65 M/hour in stampede. For GW150914, the speed on KNL is around 3.5-4M/hour compared to 12M/hour on Stampede. I also briefly looked at TimerReport but any particular thorn did not stand out. I will study it in more detail.
In general, how can I find the optimized values of 'turning knobs' (except trial and error method) and what are the constraints on them? What are the general options/parameters I can change to boost up the performance? I also had several questions about various options in machine files and about optimization and MPI in general. Can you suggest some reference where I can read more about this?
Lastly, the crashing the GW150914 in normal queue doesn't seem to be due to this reason (but I may be wrong). The error file shows segmentation fault errors. I was browsing through the past tickets and found that you had also encountered a similar segfault issue on KNL. Were you able to resolve it? I am attaching the error file, could you please look at it?
Thanks
.............................
Bhavesh Khamesra
Graduate Student
Centre of Relativistic Astrophysics
Georgia Institute of Technology
________________________________ From: schnetter@gmail.com schnetter@gmail.com on behalf of Erik Schnetter schnetter@cct.lsu.edu Sent: Wednesday, May 3, 2017 4:59:16 PM To: Khamesra, Bhavesh Cc: users@einsteintoolkit.org Subject: Re: [Users] Benchmarking
Bhavesh
To be exact, the remedy for this particular Slab error is not to use more cores, but to use more MPI processes. You can keep the number of cores constant if you reduce the number of OpenMP threads per MPI process.
Given that you are benchmarking, you should anyway experiment with these parameters, as performance can crucially depend on them. Usually, using fewer threads and more processes is more efficient for small core counts.
Finally, only comparing the overall run time is not sufficient to make a statement about performance. Each run has several "tuning knobs", and choosing the right values for these is important to achieve good performance. Using the default settings will often lead to quite poor performance. Cactus timer output as well as experience with performing runs on HPC systems is indispensable to get good performance.
-erik
On Tue, May 2, 2017 at 5:09 PM, Khamesra, Bhavesh <bhaveshkhamesra@gatech.edumailto:bhaveshkhamesra@gatech.edu> wrote:
Hi, I have sent the pull request with the optionlist for Stampede - KNL on Bitbucket simfactory repo. I have tested this with a couple of thornlists including the einsteintoolkit.thhttp://einsteintoolkit.th and GW150914.th. This is still in experimental stage and so would be great if someone could also test it.
Working on benchmarking the performance on Stampede KNL, I was able to do some test runs using the GW150914 simulation. However, I have been running into some issues with it.
1. I tried running QC0 simulation on both Stampede SandyBridge and KNL. While it runs fine on Stampede but it crashes on KNL with this error -
while executing schedule bin BoundaryConditions, routine RotatingSymmetry180::Rot180_ApplyBC in thorn RotatingSymmetry180, file /work/04082/tg833814/Cactus_ETK_dev/arrangements/CactusNumerical/RotatingSymmetry180/src/rotatingsymmetry180.c:460: -> TAT/Slab can only be used if there is a single local component per MPI process TACC: MPI job exited with code: 134 I looked up at previous tickets and found the solution to increase the number of cores. But if the same simulation can be run on stampede on 64 cores, why does it require higher number of cores on KNL? Or is it some other issue?
2. I was able to run GW150914 on development queue (68 cores) and the speeds on Stampede were around 12.9M while that on KNL goes around 2.4M. To understand the reason for such small speeds, I tried running this on higher number of cores on Stampede (128) and it runs at speed of around 20.9M (tested the run for 12 hours). However, on doing the same in normal queue in KNL, the simulation crashes after a couple of iterations on KNL with some segmentation fault error. Also, before crashing, the speed on KNL is around 4.2M. I have attached the error file of the simulation.
Could someone please look at this? Let me know if you need any other information.
Thanks
.............................
Bhavesh Khamesra
Graduate Student
Centre of Relativistic Astrophysics
Georgia Institute of Technology
_______________________________________________ Users mailing list Users@einsteintoolkit.orgmailto:Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Erik Schnetter <schnetter@cct.lsu.edumailto:schnetter@cct.lsu.edu> http://www.perimeterinstitute.ca/personal/eschnetter/
On Fri, May 5, 2017 at 11:37 AM, Khamesra, Bhavesh < bhaveshkhamesra@gatech.edu> wrote:
Hi Erik, Thanks for the reply. I tried playing with num-threads options
in machine files and was able to run the QC0 on development node. Reducing the num-threads to 17 keeping number of cores to 64
Bhavesh
This combination doesn't make sense; the number of cores needs to be a multiple of the number of threads. I would try 68 cores and 4 threads; as I mentioned, using fewer threads is currently more efficient.
showed some increase in the speed but it is still quite low - around
13-16M/hour compared the 55-65 M/hour in stampede. For GW150914, the speed on KNL is around 3.5-4M/hour compared to 12M/hour on Stampede. I also briefly looked at TimerReport but any particular thorn did not stand out. I will study it in more detail.
When you measured these speeds, how many nodes did you use each time?
Yes, you will need to look at timer output to find out e.g. whether the time is spend doing I/O.
You can also look at how much memory is used per core. As a general rule, using more memory per core is more efficient. If you are using only a very small fraction of the available memory (<10%), or if there are many more ghost points than interior points, then your setup is likely inefficient.
In general, how can I find the optimized values of 'turning knobs'
(except trial and error method) and what are the constraints on them? What are the general options/parameters I can change to boost up the performance? I also had several questions about various options in machine files and about optimization and MPI in general. Can you suggest some reference where I can read more about this?
That is a very good question. I do not have a good answer to it. (This is why it is a good question.) This information is, unfortunately, only passed on from grad student to grad student (or from postdoc to postdoc). There really should be a tutorial and some larger documentation that addresses this.
Before you can optimize the values of the knobs, you need to know what knobs there actually are.
The best offer I can make is to ask many questions while keeping notes, or to visit an experienced Einstein Toolkit user and camp out near their desk asking many questions. You can also ask people for their tuned parameter files and compare, and also ask people about their run times and speeds so that you have a basis for the comparison. (You are already doing this.)
There are three kinds of knobs that influence performance - knobs that change the physics; usually, when you write a paper, you want to keep these fixed - knobs that change the numerical approximate, such as the resolution or grid structure - knobs that change how the simulation runs, such as the number of cores, threads, or how often to do output
When running a simulation, you want to begin by tuning the third kind of knob. When you've found some reasonable optimum, you can move on to the second kind and experiment what resolution etc. you need to solve a particular physics question.
Lastly, the crashing the GW150914 in normal queue doesn't seem to be due
to this reason (but I may be wrong). The error file shows segmentation fault errors. I was browsing through the past tickets and found that you had also encountered a similar segfault issue on KNL. Were you able to resolve it? I am attaching the error file, could you please look at it?
I have heard of such a segfault before. I assume it is caused by using too many processes or too many threads for a particular resolution. I have not yet reproduced it, and I don't know what causes it. It would be helpful if you could produce a stack backtrace or similar. On the other hand, if this segfault only appears for very inefficient configurations, then there is no urgent need to debug this, as people won't be interested in using such configurations anyway.
All the best. Please keep asking.
-erik
Thanks
.............................
Bhavesh Khamesra
Graduate Student
Centre of Relativistic Astrophysics
Georgia Institute of Technology
From: schnetter@gmail.com schnetter@gmail.com on behalf of Erik
Schnetter schnetter@cct.lsu.edu
Sent: Wednesday, May 3, 2017 4:59:16 PM To: Khamesra, Bhavesh Cc: users@einsteintoolkit.org Subject: Re: [Users] Benchmarking
Bhavesh
To be exact, the remedy for this particular Slab error is not to use more
cores, but to use more MPI processes. You can keep the number of cores constant if you reduce the number of OpenMP threads per MPI process.
Given that you are benchmarking, you should anyway experiment with these
parameters, as performance can crucially depend on them. Usually, using fewer threads and more processes is more efficient for small core counts.
Finally, only comparing the overall run time is not sufficient to make a
statement about performance. Each run has several "tuning knobs", and choosing the right values for these is important to achieve good performance. Using the default settings will often lead to quite poor performance. Cactus timer output as well as experience with performing runs on HPC systems is indispensable to get good performance.
-erik
On Tue, May 2, 2017 at 5:09 PM, Khamesra, Bhavesh <
bhaveshkhamesra@gatech.edu> wrote:
Hi, I have sent the pull request with the optionlist for Stampede - KNL
on Bitbucket simfactory repo. I have tested this with a couple of thornlists including the einsteintoolkit.th and GW150914.th. This is still in experimental stage and so would be great if someone could also test it.
Working on benchmarking the performance on Stampede KNL, I was able to
do some test runs using the GW150914 simulation. However, I have been running into some issues with it.
- I tried running QC0 simulation on both Stampede SandyBridge and KNL.
While it runs fine on Stampede but it crashes on KNL with this error -
while executing schedule bin BoundaryConditions, routine
RotatingSymmetry180::Rot180_ApplyBC in thorn RotatingSymmetry180, file /work/04082/tg833814/Cactus_ETK_dev/arrangements/CactusNumerical/RotatingSymmetry180/src/rotatingsymmetry180.c:460:
-> TAT/Slab can only be used if there is a single local component per
MPI process
TACC: MPI job exited with code: 134 I looked up at previous tickets and found the solution to increase the
number of cores. But if the same simulation can be run on stampede on 64 cores, why does it require higher number of cores on KNL? Or is it some other issue?
- I was able to run GW150914 on development queue (68 cores) and the
speeds on Stampede were around 12.9M while that on KNL goes around 2.4M. To understand the reason for such small speeds, I tried running this on higher number of cores on Stampede (128) and it runs at speed of around 20.9M (tested the run for 12 hours). However, on doing the same in normal queue in KNL, the simulation crashes after a couple of iterations on KNL with some segmentation fault error. Also, before crashing, the speed on KNL is around 4.2M. I have attached the error file of the simulation.
Could someone please look at this? Let me know if you need any other
information.
Thanks
.............................
Bhavesh Khamesra
Graduate Student
Centre of Relativistic Astrophysics
Georgia Institute of Technology
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
Hi Erik,
This combination doesn't make sense; the number of cores needs to be a multiple of the number of threads. I would try 68 cores and 4 threads; as I mentioned, using fewer threads is currently more efficient. Sorry my bad, I was using 17 threads with 68 cores. I will reduce number of threads to 4. In general, how is the max and min number of threads decided (apart from the constraint you already mentioned). I was having this doubt while setting max-num-threads.
When you measured these speeds, how many nodes did you use each time?
Just one node (test done on development queue)
I have heard of such a segfault before. I assume it is caused by using too many processes or too many threads for a particular resolution. I have not yet reproduced it, and I don't know what causes it. It would be helpful if you could produce a stack backtrace or similar. On the other hand, if this segfault only appears for very inefficient configurations, then there is no urgent need to debug this, as people won't be interested in using such configurations anyway. I do not think that it is caused by using too many processors. I tried it with varying number of cores from 68 (insufficient) to 204 but get the same error. Also, I am getting this error with all my BBH parameter files I have tried, so its kind of important. Is there a way to produce backtrace file using simfactory?
Thanks
.............................
Bhavesh Khamesra
Graduate Student
Centre of Relativistic Astrophysics
Georgia Institute of Technology
________________________________ From: schnetter@gmail.com schnetter@gmail.com on behalf of Erik Schnetter schnetter@cct.lsu.edu Sent: Friday, May 5, 2017 11:52 AM To: Khamesra, Bhavesh Cc: users@einsteintoolkit.org Subject: Re: [Users] Benchmarking
On Fri, May 5, 2017 at 11:37 AM, Khamesra, Bhavesh <bhaveshkhamesra@gatech.edumailto:bhaveshkhamesra@gatech.edu> wrote:
Hi Erik, Thanks for the reply. I tried playing with num-threads options in machine files and was able to run the QC0 on development node. Reducing the num-threads to 17 keeping number of cores to 64
Bhavesh
This combination doesn't make sense; the number of cores needs to be a multiple of the number of threads. I would try 68 cores and 4 threads; as I mentioned, using fewer threads is currently more efficient.
showed some increase in the speed but it is still quite low - around 13-16M/hour compared the 55-65 M/hour in stampede. For GW150914, the speed on KNL is around 3.5-4M/hour compared to 12M/hour on Stampede. I also briefly looked at TimerReport but any particular thorn did not stand out. I will study it in more detail.
When you measured these speeds, how many nodes did you use each time?
Yes, you will need to look at timer output to find out e.g. whether the time is spend doing I/O.
You can also look at how much memory is used per core. As a general rule, using more memory per core is more efficient. If you are using only a very small fraction of the available memory (<10%), or if there are many more ghost points than interior points, then your setup is likely inefficient.
In general, how can I find the optimized values of 'turning knobs' (except trial and error method) and what are the constraints on them? What are the general options/parameters I can change to boost up the performance? I also had several questions about various options in machine files and about optimization and MPI in general. Can you suggest some reference where I can read more about this?
That is a very good question. I do not have a good answer to it. (This is why it is a good question.) This information is, unfortunately, only passed on from grad student to grad student (or from postdoc to postdoc). There really should be a tutorial and some larger documentation that addresses this.
Before you can optimize the values of the knobs, you need to know what knobs there actually are.
The best offer I can make is to ask many questions while keeping notes, or to visit an experienced Einstein Toolkit user and camp out near their desk asking many questions. You can also ask people for their tuned parameter files and compare, and also ask people about their run times and speeds so that you have a basis for the comparison. (You are already doing this.)
There are three kinds of knobs that influence performance - knobs that change the physics; usually, when you write a paper, you want to keep these fixed - knobs that change the numerical approximate, such as the resolution or grid structure - knobs that change how the simulation runs, such as the number of cores, threads, or how often to do output
When running a simulation, you want to begin by tuning the third kind of knob. When you've found some reasonable optimum, you can move on to the second kind and experiment what resolution etc. you need to solve a particular physics question.
Lastly, the crashing the GW150914 in normal queue doesn't seem to be due to this reason (but I may be wrong). The error file shows segmentation fault errors. I was browsing through the past tickets and found that you had also encountered a similar segfault issue on KNL. Were you able to resolve it? I am attaching the error file, could you please look at it?
I have heard of such a segfault before. I assume it is caused by using too many processes or too many threads for a particular resolution. I have not yet reproduced it, and I don't know what causes it. It would be helpful if you could produce a stack backtrace or similar. On the other hand, if this segfault only appears for very inefficient configurations, then there is no urgent need to debug this, as people won't be interested in using such configurations anyway.
All the best. Please keep asking.
-erik
Thanks
.............................
Bhavesh Khamesra
Graduate Student
Centre of Relativistic Astrophysics
Georgia Institute of Technology
From: schnetter@gmail.commailto:schnetter@gmail.com <schnetter@gmail.commailto:schnetter@gmail.com> on behalf of Erik Schnetter <schnetter@cct.lsu.edumailto:schnetter@cct.lsu.edu> Sent: Wednesday, May 3, 2017 4:59:16 PM To: Khamesra, Bhavesh Cc: users@einsteintoolkit.orgmailto:users@einsteintoolkit.org Subject: Re: [Users] Benchmarking
Bhavesh
To be exact, the remedy for this particular Slab error is not to use more cores, but to use more MPI processes. You can keep the number of cores constant if you reduce the number of OpenMP threads per MPI process.
Given that you are benchmarking, you should anyway experiment with these parameters, as performance can crucially depend on them. Usually, using fewer threads and more processes is more efficient for small core counts.
Finally, only comparing the overall run time is not sufficient to make a statement about performance. Each run has several "tuning knobs", and choosing the right values for these is important to achieve good performance. Using the default settings will often lead to quite poor performance. Cactus timer output as well as experience with performing runs on HPC systems is indispensable to get good performance.
-erik
On Tue, May 2, 2017 at 5:09 PM, Khamesra, Bhavesh <bhaveshkhamesra@gatech.edumailto:bhaveshkhamesra@gatech.edu> wrote:
Hi, I have sent the pull request with the optionlist for Stampede - KNL on Bitbucket simfactory repo. I have tested this with a couple of thornlists including the einsteintoolkit.thhttp://einsteintoolkit.th and GW150914.th. This is still in experimental stage and so would be great if someone could also test it.
Working on benchmarking the performance on Stampede KNL, I was able to do some test runs using the GW150914 simulation. However, I have been running into some issues with it.
- I tried running QC0 simulation on both Stampede SandyBridge and KNL. While it runs fine on Stampede but it crashes on KNL with this error -
while executing schedule bin BoundaryConditions, routine RotatingSymmetry180::Rot180_ApplyBC in thorn RotatingSymmetry180, file /work/04082/tg833814/Cactus_ETK_dev/arrangements/CactusNumerical/RotatingSymmetry180/src/rotatingsymmetry180.c:460: -> TAT/Slab can only be used if there is a single local component per MPI process TACC: MPI job exited with code: 134 I looked up at previous tickets and found the solution to increase the number of cores. But if the same simulation can be run on stampede on 64 cores, why does it require higher number of cores on KNL? Or is it some other issue?
- I was able to run GW150914 on development queue (68 cores) and the speeds on Stampede were around 12.9M while that on KNL goes around 2.4M. To understand the reason for such small speeds, I tried running this on higher number of cores on Stampede (128) and it runs at speed of around 20.9M (tested the run for 12 hours). However, on doing the same in normal queue in KNL, the simulation crashes after a couple of iterations on KNL with some segmentation fault error. Also, before crashing, the speed on KNL is around 4.2M. I have attached the error file of the simulation.
Could someone please look at this? Let me know if you need any other information.
Thanks
.............................
Bhavesh Khamesra
Graduate Student
Centre of Relativistic Astrophysics
Georgia Institute of Technology
Users mailing list Users@einsteintoolkit.orgmailto:Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Erik Schnetter <schnetter@cct.lsu.edumailto:schnetter@cct.lsu.edu> http://www.perimeterinstitute.ca/personal/eschnetter/
-- Erik Schnetter <schnetter@cct.lsu.edumailto:schnetter@cct.lsu.edu> http://www.perimeterinstitute.ca/personal/eschnetter/
On 5 May 2017, at 18:43, Khamesra, Bhavesh bhaveshkhamesra@gatech.edu wrote:
I have heard of such a segfault before. I assume it is caused by using too many processes or too many threads for a particular resolution. I have not yet reproduced it, and I don't know what causes it. It would be helpful if you could produce a stack backtrace or similar. On the other hand, if this segfault only appears for very inefficient configurations, then there is no urgent need to debug this, as people won't be interested in using such configurations anyway. I do not think that it is caused by using too many processors. I tried it with varying number of cores from 68 (insufficient) to 204 but get the same error. Also, I am getting this error with all my BBH parameter files I have tried, so its kind of important. Is there a way to produce backtrace file using simfactory?
Carpet should write a backtrace file into the output directory if there is a segfault. Do you see backtrace files in
sim/output-0000/sim/backtrace.*.txt
?
Hi Ian, No I do not find any backtrace files in the output directory.
.............................
Bhavesh Khamesra
Graduate Student
Centre of Relativistic Astrophysics
Georgia Institute of Technology
________________________________ From: Ian Hinder ian.hinder@aei.mpg.de Sent: Friday, May 5, 2017 1:13:38 PM To: Khamesra, Bhavesh Cc: Erik Schnetter; Einstein Toolkit Users Subject: Re: [Users] Benchmarking
On 5 May 2017, at 18:43, Khamesra, Bhavesh <bhaveshkhamesra@gatech.edumailto:bhaveshkhamesra@gatech.edu> wrote:
I have heard of such a segfault before. I assume it is caused by using too many processes or too many threads for a particular resolution. I have not yet reproduced it, and I don't know what causes it. It would be helpful if you could produce a stack backtrace or similar. On the other hand, if this segfault only appears for very inefficient configurations, then there is no urgent need to debug this, as people won't be interested in using such configurations anyway. I do not think that it is caused by using too many processors. I tried it with varying number of cores from 68 (insufficient) to 204 but get the same error. Also, I am getting this error with all my BBH parameter files I have tried, so its kind of important. Is there a way to produce backtrace file using simfactory?
Carpet should write a backtrace file into the output directory if there is a segfault. Do you see backtrace files in
sim/output-0000/sim/backtrace.*.txt
?
-- Ian Hinder http://members.aei.mpg.de/ianhin
On 5 May 2017, at 20:38, Khamesra, Bhavesh bhaveshkhamesra@gatech.edu wrote:
Hi Ian, No I do not find any backtrace files in the output directory.
How do you know that there is a segfault? Can you post the error message? It's possible that the segfault is not in Cactus - e.g. it could be from mpirun or some other part of the system. If Cactus aborts with a segfault, and the run is using Carpet as the driver (i.e. the Carpet thorn is active in the parameter file), then it should write a backtrace file.
............................. Bhavesh Khamesra Graduate Student Centre of Relativistic Astrophysics Georgia Institute of Technology From: Ian Hinder ian.hinder@aei.mpg.de Sent: Friday, May 5, 2017 1:13:38 PM To: Khamesra, Bhavesh Cc: Erik Schnetter; Einstein Toolkit Users Subject: Re: [Users] Benchmarking
On 5 May 2017, at 18:43, Khamesra, Bhavesh bhaveshkhamesra@gatech.edu wrote:
I have heard of such a segfault before. I assume it is caused by using too many processes or too many threads for a particular resolution. I have not yet reproduced it, and I don't know what causes it. It would be helpful if you could produce a stack backtrace or similar. On the other hand, if this segfault only appears for very inefficient configurations, then there is no urgent need to debug this, as people won't be interested in using such configurations anyway. I do not think that it is caused by using too many processors. I tried it with varying number of cores from 68 (insufficient) to 204 but get the same error. Also, I am getting this error with all my BBH parameter files I have tried, so its kind of important. Is there a way to produce backtrace file using simfactory?
Carpet should write a backtrace file into the output directory if there is a segfault. Do you see backtrace files in
sim/output-0000/sim/backtrace.*.txt
?
-- Ian Hinder http://members.aei.mpg.de/ianhin
Hi Ian, The segfault was mentioned in the error file, you can find the complete error in it (Attached here).
It's possible that the segfault is not in Cactus
Yes, I completely agree. The error does not refer to any Cactus file and hence I wasn't sure what is the cause of it.
Thanks,
.............................
Bhavesh Khamesra
Graduate Student
Centre of Relativistic Astrophysics
Georgia Institute of Technology
________________________________ From: Ian Hinder ian.hinder@aei.mpg.de Sent: Saturday, May 6, 2017 3:58:46 AM To: Khamesra, Bhavesh Cc: Erik Schnetter; Einstein Toolkit Users Subject: Re: [Users] Benchmarking
On 5 May 2017, at 20:38, Khamesra, Bhavesh <bhaveshkhamesra@gatech.edumailto:bhaveshkhamesra@gatech.edu> wrote:
Hi Ian, No I do not find any backtrace files in the output directory.
How do you know that there is a segfault? Can you post the error message? It's possible that the segfault is not in Cactus - e.g. it could be from mpirun or some other part of the system. If Cactus aborts with a segfault, and the run is using Carpet as the driver (i.e. the Carpet thorn is active in the parameter file), then it should write a backtrace file.
............................. Bhavesh Khamesra Graduate Student Centre of Relativistic Astrophysics Georgia Institute of Technology ________________________________ From: Ian Hinder <ian.hinder@aei.mpg.demailto:ian.hinder@aei.mpg.de> Sent: Friday, May 5, 2017 1:13:38 PM To: Khamesra, Bhavesh Cc: Erik Schnetter; Einstein Toolkit Users Subject: Re: [Users] Benchmarking
On 5 May 2017, at 18:43, Khamesra, Bhavesh <bhaveshkhamesra@gatech.edumailto:bhaveshkhamesra@gatech.edu> wrote:
I have heard of such a segfault before. I assume it is caused by using too many processes or too many threads for a particular resolution. I have not yet reproduced it, and I don't know what causes it. It would be helpful if you could produce a stack backtrace or similar. On the other hand, if this segfault only appears for very inefficient configurations, then there is no urgent need to debug this, as people won't be interested in using such configurations anyway. I do not think that it is caused by using too many processors. I tried it with varying number of cores from 68 (insufficient) to 204 but get the same error. Also, I am getting this error with all my BBH parameter files I have tried, so its kind of important. Is there a way to produce backtrace file using simfactory?
Carpet should write a backtrace file into the output directory if there is a segfault. Do you see backtrace files in
sim/output-0000/sim/backtrace.*.txt
?
-- Ian Hinder http://members.aei.mpg.de/ianhin
-- Ian Hinder http://members.aei.mpg.de/ianhin
On 6 May 2017, at 14:28, Khamesra, Bhavesh bhaveshkhamesra@gatech.edu wrote:
Hi Ian, The segfault was mentioned in the error file, you can find the complete error in it (Attached here).
It's possible that the segfault is not in Cactus Yes, I completely agree. The error does not refer to any Cactus file and hence I wasn't sure what is the cause of it.
/usr/local/bin/tacc_affinity: line 58: 43015 Segmentation fault $@
Does Cactus finish successfully? It might be something to ask the TACC admins about, since tacc_affinity is their script.
Does Cactus finish successfully?
No, it ends with an MPI exit code 134. I will send an email to TACC people. Thanks a lot Ian.
Cheers
.............................
Bhavesh Khamesra
Graduate Student
Centre of Relativistic Astrophysics
Georgia Institute of Technology
________________________________ From: Ian Hinder ian.hinder@aei.mpg.de Sent: Sunday, May 7, 2017 3:10:41 PM To: Khamesra, Bhavesh Cc: Erik Schnetter; Einstein Toolkit Users Subject: Re: [Users] Benchmarking
On 6 May 2017, at 14:28, Khamesra, Bhavesh <bhaveshkhamesra@gatech.edumailto:bhaveshkhamesra@gatech.edu> wrote:
Hi Ian, The segfault was mentioned in the error file, you can find the complete error in it (Attached here).
It's possible that the segfault is not in Cactus Yes, I completely agree. The error does not refer to any Cactus file and hence I wasn't sure what is the cause of it.
/usr/local/bin/tacc_affinity: line 58: 43015 Segmentation fault $@
Does Cactus finish successfully? It might be something to ask the TACC admins about, since tacc_affinity is their script.
-- Ian Hinder http://members.aei.mpg.de/ianhin
users@lists.einsteintoolkit.org