Hello,
Is there documentation about performance of Cactus ETK in large machines. I have some questions regarding best performance according to initial conditions, calculation time required, etc. If there are performance plots like Flops vs. Number of nodes would help me as well.
Other main question I have is if Cactus scales in large machines.
I am using McLachlan, but any other application would give me an idea of what I should expect for my runs in big machines.
Thanks, Jose
Hi,
On Tue, Mar 20, 2012 at 05:14:38PM -0700, Jose Fiestas Iquira wrote:
Is there documentation about performance of Cactus ETK in large machines. I have some questions regarding best performance according to initial conditions, calculation time required, etc.
Performance very much depends on the specific setup. One poorly scaling function can ruin the otherwise best run.
If there are performance plots like Flops vs. Number of nodes would help me as well.
Flops are very problem-dependent. There isn't such thing as flops/s for Cactus, not even for one given machine. If we talk about the Einstein equations and a typical production run I would expect a few percent of the peak performance of any given CPU, as we are most of the time bound by memory bandwidth. Calculation time for initial data very much depends on the type of initial data. Some initial data are setup in a few seconds, some may need a day. Scaling of initial data computation also might be quite different from that of the evolution, which is why sometimes it makes sense to checkpoint right after initial data setup and restart using a different number of cores.
Other main question I have is if Cactus scales in large machines.
I am afraid that this question is too general. 'Cactus' itself doesn't even deal directly with MPI, so it would be better to ask, e.g., how well Carpet scales. And given that still quite general question it can be said that Carpet scales to at least 100k cores - however, that again very much depends on your setup. Using a couple of levels of mesh refinement bring down the scaling limit to maybe 10k cores, adding re-gridding and a few common analysis routines and we talk about 2k-4k cores practical limit. But again, these numbers very much depend on your setup. You might be able to perform much better for certain problems, and much worse for others.
I am using McLachlan, but any other application would give me an idea of what I should expect for my runs in big machines.
Do the numbers I gave help?
Frank
On Tue, Mar 20, 2012 at 10:45 PM, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi,
On Tue, Mar 20, 2012 at 05:14:38PM -0700, Jose Fiestas Iquira wrote:
Is there documentation about performance of Cactus ETK in large machines. I have some questions regarding best performance according to initial conditions, calculation time required, etc.
Performance very much depends on the specific setup. One poorly scaling function can ruin the otherwise best run.
If there are performance plots like Flops vs. Number of nodes would help me as well.
Flops are very problem-dependent. There isn't such thing as flops/s for Cactus, not even for one given machine. If we talk about the Einstein equations and a typical production run I would expect a few percent of the peak performance of any given CPU, as we are most of the time bound by memory bandwidth.
I would like to add some more numbers to Frank's description:
One some problems (e.g. evaluating the BSSN equations with a higher-order stencil), I have measured more than 20% of the theoretical peak performance. The bottleneck seem to be L1 data cache accesses, because the BSSN equation kernels require a large number of local (temporary) variables.
If you look for parallel scaling, then e.g. http://arxiv.org/abs/1111.3344 contains a scaling graph for the BSSN equations evolved with mesh refinement. This shows that, for this benchmark, the Einstein Toolkit scales well to more than 12k cores.
-erik
Hello,
I reduced the simulation time by setting Cactus::cctk_final_time = .01 in order to measure performance with CrayPat. It run only 8 iterations. I used 16 and 24 cores for testing, and obtained almost the same performance (~1310 sec. simulation time, and ~16MFlops).
It remembers me Fig.2 in the reference you sent http://arxiv.org/abs/1111.3344
which I don't really understand. I would expect shorter times with larger number of cores. Why does it not happen here?
I am using McLachlan to simulate a binary system. So, all my regards are concerning this specific application. Do you think it will scale in the sense that simulation time will be shorter, the larger of number of cores I use?
Thanks, Jose
On Wed, Mar 21, 2012 at 5:08 AM, Erik Schnetter schnetter@cct.lsu.eduwrote:
On Tue, Mar 20, 2012 at 10:45 PM, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi,
On Tue, Mar 20, 2012 at 05:14:38PM -0700, Jose Fiestas Iquira wrote:
Is there documentation about performance of Cactus ETK in large
machines. I
have some questions regarding best performance according to initial conditions, calculation time required, etc.
Performance very much depends on the specific setup. One poorly scaling function can ruin the otherwise best run.
If there are performance plots like Flops vs. Number of nodes would
help me
as well.
Flops are very problem-dependent. There isn't such thing as flops/s for Cactus, not even for one given machine. If we talk about the Einstein equations and a typical production run I would expect a few percent of the peak performance of any given CPU, as we are most of the time bound
by
memory bandwidth.
I would like to add some more numbers to Frank's description:
One some problems (e.g. evaluating the BSSN equations with a higher-order stencil), I have measured more than 20% of the theoretical peak performance. The bottleneck seem to be L1 data cache accesses, because the BSSN equation kernels require a large number of local (temporary) variables.
If you look for parallel scaling, then e.g. http://arxiv.org/abs/1111.3344 contains a scaling graph for the BSSN equations evolved with mesh refinement. This shows that, for this benchmark, the Einstein Toolkit scales well to more than 12k cores.
-erik
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
Hi Jose,
look, the Einstein Toolkit team is very happy to help new users like you to get started and sort out specific questions regarding parts of the toolkit.
What we really can't do is provide you with very basic high-performance computing training via the mailing list. This is because many if not most people on this list actually volunteer to help in their spare time and are not paid as consultants for general HPC questions. You are at Berkeley lab and there are many experts that can help you with basic HPC questions, plus there are tons of resources available on-line, that I would kindly ask you to consult first.
Regarding your scaling question:
https://support.scinet.utoronto.ca/wiki/index.php/Introduction_To_Performanc...
gives a good introduction to performance measurements. There are many more webpages like this available on the internet. The plot shown in the Einstein Toolkit paper (arXiv:1111.3344) is a weak scaling test.
Best,
- Christian Ott
On Wed, Mar 28, 2012 at 11:55:30PM -0700, Jose Fiestas Iquira wrote:
Hello,
I reduced the simulation time by setting Cactus::cctk_final_time = .01 in order to measure performance with CrayPat. It run only 8 iterations. I used 16 and 24 cores for testing, and obtained almost the same performance (~1310 sec. simulation time, and ~16MFlops).
It remembers me Fig.2 in the reference you sent http://arxiv.org/abs/1111.3344
which I don't really understand. I would expect shorter times with larger number of cores. Why does it not happen here?
I am using McLachlan to simulate a binary system. So, all my regards are concerning this specific application. Do you think it will scale in the sense that simulation time will be shorter, the larger of number of cores I use?
Thanks, Jose
On Wed, Mar 21, 2012 at 5:08 AM, Erik Schnetter schnetter@cct.lsu.eduwrote:
On Tue, Mar 20, 2012 at 10:45 PM, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi,
On Tue, Mar 20, 2012 at 05:14:38PM -0700, Jose Fiestas Iquira wrote:
Is there documentation about performance of Cactus ETK in large
machines. I
have some questions regarding best performance according to initial conditions, calculation time required, etc.
Performance very much depends on the specific setup. One poorly scaling function can ruin the otherwise best run.
If there are performance plots like Flops vs. Number of nodes would
help me
as well.
Flops are very problem-dependent. There isn't such thing as flops/s for Cactus, not even for one given machine. If we talk about the Einstein equations and a typical production run I would expect a few percent of the peak performance of any given CPU, as we are most of the time bound
by
memory bandwidth.
I would like to add some more numbers to Frank's description:
One some problems (e.g. evaluating the BSSN equations with a higher-order stencil), I have measured more than 20% of the theoretical peak performance. The bottleneck seem to be L1 data cache accesses, because the BSSN equation kernels require a large number of local (temporary) variables.
If you look for parallel scaling, then e.g. http://arxiv.org/abs/1111.3344 contains a scaling graph for the BSSN equations evolved with mesh refinement. This shows that, for this benchmark, the Einstein Toolkit scales well to more than 12k cores.
-erik
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Dear Christian, Thank you. I understand what you are saying. I am mainly asking regarding McLachlan. Sorry if it appears I want to learn about HPC from you. I am working with people in the lab, for sure. They are just not aware about Cactus and I am learning as well. I apologize for that, and will try to avoid it in the future. Sincerely, Jose
On Thu, Mar 29, 2012 at 8:01 AM, Christian D. Ott cott@tapir.caltech.eduwrote:
Hi Jose,
look, the Einstein Toolkit team is very happy to help new users like you to get started and sort out specific questions regarding parts of the toolkit.
What we really can't do is provide you with very basic high-performance computing training via the mailing list. This is because many if not most people on this list actually volunteer to help in their spare time and are not paid as consultants for general HPC questions. You are at Berkeley lab and there are many experts that can help you with basic HPC questions, plus there are tons of resources available on-line, that I would kindly ask you to consult first.
Regarding your scaling question:
https://support.scinet.utoronto.ca/wiki/index.php/Introduction_To_Performanc...
gives a good introduction to performance measurements. There are many more webpages like this available on the internet. The plot shown in the Einstein Toolkit paper (arXiv:1111.3344) is a weak scaling test.
Best,
- Christian Ott
On Wed, Mar 28, 2012 at 11:55:30PM -0700, Jose Fiestas Iquira wrote:
Hello,
I reduced the simulation time by setting Cactus::cctk_final_time = .01 in order to measure performance with CrayPat. It run only 8 iterations. I
used
16 and 24 cores for testing, and obtained almost the same performance (~1310 sec. simulation time, and ~16MFlops).
It remembers me Fig.2 in the reference you sent http://arxiv.org/abs/1111.3344
which I don't really understand. I would expect shorter times with larger number of cores. Why does it not happen here?
I am using McLachlan to simulate a binary system. So, all my regards are concerning this specific application. Do you think it will scale in the sense that simulation time will be shorter, the larger of number of
cores I
use?
Thanks, Jose
On Wed, Mar 21, 2012 at 5:08 AM, Erik Schnetter <schnetter@cct.lsu.edu wrote:
On Tue, Mar 20, 2012 at 10:45 PM, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi,
On Tue, Mar 20, 2012 at 05:14:38PM -0700, Jose Fiestas Iquira wrote:
Is there documentation about performance of Cactus ETK in large
machines. I
have some questions regarding best performance according to initial conditions, calculation time required, etc.
Performance very much depends on the specific setup. One poorly
scaling
function can ruin the otherwise best run.
If there are performance plots like Flops vs. Number of nodes would
help me
as well.
Flops are very problem-dependent. There isn't such thing as flops/s
for
Cactus, not even for one given machine. If we talk about the Einstein equations and a typical production run I would expect a few percent
of
the peak performance of any given CPU, as we are most of the time
bound
by
memory bandwidth.
I would like to add some more numbers to Frank's description:
One some problems (e.g. evaluating the BSSN equations with a higher-order stencil), I have measured more than 20% of the theoretical peak performance. The bottleneck seem to be L1 data cache accesses, because the BSSN equation kernels require a large number of local (temporary) variables.
If you look for parallel scaling, then e.g. http://arxiv.org/abs/1111.3344 contains a scaling graph for the BSSN equations evolved with mesh refinement. This shows that, for this benchmark, the Einstein Toolkit scales well to more than 12k cores.
-erik
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 29 Mar 2012, at 08:55, Jose Fiestas Iquira wrote:
Hello,
I reduced the simulation time by setting Cactus::cctk_final_time = .01 in order to measure performance with CrayPat. It run only 8 iterations. I used 16 and 24 cores for testing, and obtained almost the same performance (~1310 sec. simulation time, and ~16MFlops).
It remembers me Fig.2 in the reference you sent http://arxiv.org/abs/1111.3344
which I don't really understand. I would expect shorter times with larger number of cores. Why does it not happen here?
Hi Jose,
You will need to investigate this yourself by looking at the timings from the different parts of the code. Look at the timer parameters in the Carpet and TimerReport thorns. I recommend Carpet::output_timer_tree_every = 32 or so, though this only gives you timings from process 0. The code is very complex and there could be many factors inhibiting scaling if you have not been very careful to set things up in a good way. One possibility is that the initial data solver might take a significant percentage of your run time, depending on the parameters you have set, and this solver is not parallelised. You should look at the time spent in Evolve, not the time spent in Initialise.
I am using McLachlan to simulate a binary system. So, all my regards are concerning this specific application. Do you think it will scale in the sense that simulation time will be shorter, the larger of number of cores I use?
In some situations yes, in others no. There is no general answer to this type of question. Both very small and very large problems will likely scale badly. There is certainly a region in-between where you expect to see good scaling with Cactus/Carpet/McLachlan.
Thanks, Jose
On Wed, Mar 21, 2012 at 5:08 AM, Erik Schnetter schnetter@cct.lsu.edu wrote: On Tue, Mar 20, 2012 at 10:45 PM, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi,
On Tue, Mar 20, 2012 at 05:14:38PM -0700, Jose Fiestas Iquira wrote:
Is there documentation about performance of Cactus ETK in large machines. I have some questions regarding best performance according to initial conditions, calculation time required, etc.
Performance very much depends on the specific setup. One poorly scaling function can ruin the otherwise best run.
If there are performance plots like Flops vs. Number of nodes would help me as well.
Flops are very problem-dependent. There isn't such thing as flops/s for Cactus, not even for one given machine. If we talk about the Einstein equations and a typical production run I would expect a few percent of the peak performance of any given CPU, as we are most of the time bound by memory bandwidth.
I would like to add some more numbers to Frank's description:
One some problems (e.g. evaluating the BSSN equations with a higher-order stencil), I have measured more than 20% of the theoretical peak performance. The bottleneck seem to be L1 data cache accesses, because the BSSN equation kernels require a large number of local (temporary) variables.
If you look for parallel scaling, then e.g. http://arxiv.org/abs/1111.3344 contains a scaling graph for the BSSN equations evolved with mesh refinement. This shows that, for this benchmark, the Einstein Toolkit scales well to more than 12k cores.
-erik
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 21 Mar 2012, at 03:45, Frank Loeffler wrote:
Hi,
On Tue, Mar 20, 2012 at 05:14:38PM -0700, Jose Fiestas Iquira wrote:
Is there documentation about performance of Cactus ETK in large machines. I have some questions regarding best performance according to initial conditions, calculation time required, etc.
Performance very much depends on the specific setup. One poorly scaling function can ruin the otherwise best run.
If there are performance plots like Flops vs. Number of nodes would help me as well.
Flops are very problem-dependent. There isn't such thing as flops/s for Cactus, not even for one given machine. If we talk about the Einstein equations and a typical production run I would expect a few percent of the peak performance of any given CPU, as we are most of the time bound by memory bandwidth.
Hi Frank,
Do you have some tests/numbers that support his? My recollection is that we get ~30% of the peak performance, though this wouldn't have been in production runs. I'm also a bit surprised by the memory bandwidth statement, though it is certainly possible.
Calculation time for initial data very much depends on the type of initial data. Some initial data are setup in a few seconds, some may need a day. Scaling of initial data computation also might be quite different from that of the evolution, which is why sometimes it makes sense to checkpoint right after initial data setup and restart using a different number of cores.
Other main question I have is if Cactus scales in large machines.
I am afraid that this question is too general. 'Cactus' itself doesn't even deal directly with MPI, so it would be better to ask, e.g., how well Carpet scales. And given that still quite general question it can be said that Carpet scales to at least 100k cores - however, that again very much depends on your setup. Using a couple of levels of mesh refinement bring down the scaling limit to maybe 10k cores, adding re-gridding and a few common analysis routines and we talk about 2k-4k cores practical limit. But again, these numbers very much depend on your setup. You might be able to perform much better for certain problems, and much worse for others.
I was able to scale the first few iterations of a production BBH simulation up to 2400 cores of our cluster (strong scaling) without losing too much performance. If you need firm numbers I can look them up.
I am using McLachlan, but any other application would give me an idea of what I should expect for my runs in big machines.
Do the numbers I gave help?
You can refer to the report of the XiRel project, which investigated scaling of our production Einstein codes up to large numbers of cores:
Jian Tao, Gabrielle Allen, Ian Hinder, Erik Schnetter, and Yosef Zlochower. XiRel: Standard benchmarks for numerical relativity codes using Cactus and Carpet. Technical Report CCT-TR-2008-5, Louisiana State University, 2008.
Unfortunately, the web is full of dead links to this project, and I can't find anything on the CCT web pages which works.
Frank, do you know what the following links have changed into?
http://www.cct.lsu.edu/xirel/ http://www.cct.lsu.edu/CCT-TR/CCT-TR-2008-5
On Wed, Mar 21, 2012 at 01:15:05PM +0100, Ian Hinder wrote:
Frank, do you know what the following links have changed into?
http://www.cct.lsu.edu/xirel/ http://www.cct.lsu.edu/CCT-TR/CCT-TR-2008-5
Good question, I made an inquiry.
Frank
Hello,
Thanks for the information. I had a look at the benchmarks and profiling links on the Cactus website, and would like to find out the way to calculate Flops performed by a Cactus application. Are they some numbers published? In one of the papers you sent to me I found timings of McLachlan (which is my application), but I could not find Flops per simulation.
It seems to me I could be able to run my application using .par-files prepared for benchmarking? Please correct me if I am wrong.
My goal is to have an idea of the Flops expected by McLachlan runs on larger machines (few thousands or cores).
Thanks, Jose
2012/3/21 Frank Loeffler knarf@cct.lsu.edu
On Wed, Mar 21, 2012 at 01:15:05PM +0100, Ian Hinder wrote:
Frank, do you know what the following links have changed into?
http://www.cct.lsu.edu/xirel/ http://www.cct.lsu.edu/CCT-TR/CCT-TR-2008-5
Good question, I made an inquiry.
Frank
-----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.10 (GNU/Linux)
iQIcBAEBCAAGBQJPaelqAAoJEOkzpip+I59kRocP/RZl0R9MEiM5r/HfB+JZkwXs cA3eXKJeqiHtzsmqqruH/AI7zcR8zpq+BsrWHmtJQecqp28JqoV1+G7V8Z7cB/os 4pI0KFQQ6npRmnWytBSLrextanvLooqFgEB62S78MsyucTgzJwVX1AOkDuYQpHPO 3V6LQO7Jmcj5m2OoVr74mI/IFDermdwE/84dyjbz5tRrKPTug7qjfjFrfKT9yysK unA6oKl9zEdm1KZvDKzPTPh6hCcI1NNC+uRoNMnBu3XQprgNvZ/N8iDomrodCjhO iqKy5DVBhZVJhh//dWYuTzRu/l0MPMv4WML0OJmeLgbYerW2m+17lwAgumtccqwk OOuGTUqVfF9VuHJQ17jZ4jVyAjl9pIHENJCDZjmwNBDONinUC3Igjza3YlSHOCzU 1QUU1eomrSvA1oKThfaTKj+vpd+IjncucmQAtVW5rsiMjKQoOMajyR13kMJ6e8Y6 uooQr8wXJBJqiN48QBS3XvGTixvAeXmT+bQtBpocyr6vN4vEPBfwWymTn2sfd/XA vd/7Rps/IH9x3bbkk3eqCHiHrxvqPBfNdm/KbI20SZUT88XbFawuL7l1X3UgKCLv fjRNPFxtHQsz3CYk/z6YllKgf6xp34XYshtGnl6LE9lS2If8GOXIMR+MdgGJndre FXx7kXx6Oj4TLYsqOE0f =vnYk -----END PGP SIGNATURE-----
Dear Jose,
I think the best is you use a performance checking tool coming with your computer. For example PAPI or
CrayPat (on your lbl machine: http://www.nersc.gov/users/software/debugging-and-profiling/craypat/). You need to access hardware counters for getting the flops; this is not a software feature.Counting flops is so last century, though, and nobody talks about this these days. I am usually just interested in getting the shortest possible walltime for a given problem I want to solve with my Uschi evolution code in the Cactus framework.
Uschi
---------------- Uschi T. Gamma Assistant to the SC, TAPIR, Caltech
________________________________ From: Jose Fiestas Iquira jafiestas@lbl.gov To: Ian Hinder ian.hinder@aei.mpg.de; users@einsteintoolkit.org Sent: Monday, March 26, 2012 8:58 PM Subject: Re: [Users] cactus performance
Hello,
Thanks for the information. I had a look at the benchmarks and profiling links on the Cactus website, and would like to find out the way to calculate Flops performed by a Cactus application. Are they some numbers published? In one of the papers you sent to me I found timings of McLachlan (which is my application), but I could not find Flops per simulation.
It seems to me I could be able to run my application using .par-files prepared for benchmarking? Please correct me if I am wrong.
My goal is to have an idea of the Flops expected by McLachlan runs on larger machines (few thousands or cores).
Thanks, Jose
2012/3/21 Frank Loeffler knarf@cct.lsu.edu
On Wed, Mar 21, 2012 at 01:15:05PM +0100, Ian Hinder wrote:
Frank, do you know what the following links have changed into?
http://www.cct.lsu.edu/xirel/ http://www.cct.lsu.edu/CCT-TR/CCT-TR-2008-5
Good question, I made an inquiry.
Frank
-----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.10 (GNU/Linux)
iQIcBAEBCAAGBQJPaelqAAoJEOkzpip+I59kRocP/RZl0R9MEiM5r/HfB+JZkwXs cA3eXKJeqiHtzsmqqruH/AI7zcR8zpq+BsrWHmtJQecqp28JqoV1+G7V8Z7cB/os 4pI0KFQQ6npRmnWytBSLrextanvLooqFgEB62S78MsyucTgzJwVX1AOkDuYQpHPO 3V6LQO7Jmcj5m2OoVr74mI/IFDermdwE/84dyjbz5tRrKPTug7qjfjFrfKT9yysK unA6oKl9zEdm1KZvDKzPTPh6hCcI1NNC+uRoNMnBu3XQprgNvZ/N8iDomrodCjhO iqKy5DVBhZVJhh//dWYuTzRu/l0MPMv4WML0OJmeLgbYerW2m+17lwAgumtccqwk OOuGTUqVfF9VuHJQ17jZ4jVyAjl9pIHENJCDZjmwNBDONinUC3Igjza3YlSHOCzU 1QUU1eomrSvA1oKThfaTKj+vpd+IjncucmQAtVW5rsiMjKQoOMajyR13kMJ6e8Y6 uooQr8wXJBJqiN48QBS3XvGTixvAeXmT+bQtBpocyr6vN4vEPBfwWymTn2sfd/XA vd/7Rps/IH9x3bbkk3eqCHiHrxvqPBfNdm/KbI20SZUT88XbFawuL7l1X3UgKCLv fjRNPFxtHQsz3CYk/z6YllKgf6xp34XYshtGnl6LE9lS2If8GOXIMR+MdgGJndre FXx7kXx6Oj4TLYsqOE0f =vnYk -----END PGP SIGNATURE-----
_______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Dear Uschi,
Thanks. I am right now trying CrayPat. I learned about it some days ago and tried my own N-cody code with CrayPat. I am using now the McLachlan Cactus thorn for my applications.
I am just not sure about the way to compile the code before running CrayPat. I was using simfactory for compilation and now I am using gmake directly. Since compilation takes some time, I am waiting for it to finish. Let see if it works.
The reason I am interested in Flops is because I want to know if my Cactus application will scale together with my own N-body runs in large machines and reach Petascale.
Btw, did you try CrayPat with Cactus? You use gmake for compiling?
Best, Jose
On Mon, Mar 26, 2012 at 10:20 PM, Ursula Gamma uschigamma@yahoo.com wrote:
Dear Jose,
I think the best is you use a performance checking tool coming with your computer. For example PAPI or CrayPat (on your lbl machine: http://www.nersc.gov/users/software/debugging-and-profiling/craypat/). You need to access hardware counters for getting the flops; this is not a software feature. Counting flops is so last century, though, and nobody talks about this these days. I am usually just interested in getting the shortest possible walltime for a given problem I want to solve with my Uschi evolution code in the Cactus framework.
Uschi
Uschi T. Gamma Assistant to the SC, TAPIR, Caltech
*From:* Jose Fiestas Iquira jafiestas@lbl.gov *To:* Ian Hinder ian.hinder@aei.mpg.de; users@einsteintoolkit.org *Sent:* Monday, March 26, 2012 8:58 PM
*Subject:* Re: [Users] cactus performance
Hello,
Thanks for the information. I had a look at the benchmarks and profiling links on the Cactus website, and would like to find out the way to calculate Flops performed by a Cactus application. Are they some numbers published? In one of the papers you sent to me I found timings of McLachlan (which is my application), but I could not find Flops per simulation.
It seems to me I could be able to run my application using .par-files prepared for benchmarking? Please correct me if I am wrong.
My goal is to have an idea of the Flops expected by McLachlan runs on larger machines (few thousands or cores).
Thanks, Jose
2012/3/21 Frank Loeffler knarf@cct.lsu.edu
On Wed, Mar 21, 2012 at 01:15:05PM +0100, Ian Hinder wrote:
Frank, do you know what the following links have changed into?
http://www.cct.lsu.edu/xirel/ http://www.cct.lsu.edu/CCT-TR/CCT-TR-2008-5
Good question, I made an inquiry.
Frank
-----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.10 (GNU/Linux)
iQIcBAEBCAAGBQJPaelqAAoJEOkzpip+I59kRocP/RZl0R9MEiM5r/HfB+JZkwXs cA3eXKJeqiHtzsmqqruH/AI7zcR8zpq+BsrWHmtJQecqp28JqoV1+G7V8Z7cB/os 4pI0KFQQ6npRmnWytBSLrextanvLooqFgEB62S78MsyucTgzJwVX1AOkDuYQpHPO 3V6LQO7Jmcj5m2OoVr74mI/IFDermdwE/84dyjbz5tRrKPTug7qjfjFrfKT9yysK unA6oKl9zEdm1KZvDKzPTPh6hCcI1NNC+uRoNMnBu3XQprgNvZ/N8iDomrodCjhO iqKy5DVBhZVJhh//dWYuTzRu/l0MPMv4WML0OJmeLgbYerW2m+17lwAgumtccqwk OOuGTUqVfF9VuHJQ17jZ4jVyAjl9pIHENJCDZjmwNBDONinUC3Igjza3YlSHOCzU 1QUU1eomrSvA1oKThfaTKj+vpd+IjncucmQAtVW5rsiMjKQoOMajyR13kMJ6e8Y6 uooQr8wXJBJqiN48QBS3XvGTixvAeXmT+bQtBpocyr6vN4vEPBfwWymTn2sfd/XA vd/7Rps/IH9x3bbkk3eqCHiHrxvqPBfNdm/KbI20SZUT88XbFawuL7l1X3UgKCLv fjRNPFxtHQsz3CYk/z6YllKgf6xp34XYshtGnl6LE9lS2If8GOXIMR+MdgGJndre FXx7kXx6Oj4TLYsqOE0f =vnYk -----END PGP SIGNATURE-----
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Hello, Regarding performance. I am willing to run McLachlan shortly using CrayPat and I am setting the simulation time here in par/qc0-mclachlan.par, like this:
Cactus::terminate = "time" Cactus::cctk_final_time = .1
usually this time was 100.
Am I missing something else in setting the time for a short simulation? I could not find it in the documentation.
Thanks, Jose
On Mon, Mar 26, 2012 at 10:44 PM, Jose Fiestas Iquira jafiestas@lbl.govwrote:
Dear Uschi,
Thanks. I am right now trying CrayPat. I learned about it some days ago and tried my own N-cody code with CrayPat. I am using now the McLachlan Cactus thorn for my applications.
I am just not sure about the way to compile the code before running CrayPat. I was using simfactory for compilation and now I am using gmake directly. Since compilation takes some time, I am waiting for it to finish. Let see if it works.
The reason I am interested in Flops is because I want to know if my Cactus application will scale together with my own N-body runs in large machines and reach Petascale.
Btw, did you try CrayPat with Cactus? You use gmake for compiling?
Best, Jose
On Mon, Mar 26, 2012 at 10:20 PM, Ursula Gamma uschigamma@yahoo.comwrote:
Dear Jose,
I think the best is you use a performance checking tool coming with your computer. For example PAPI or CrayPat (on your lbl machine: http://www.nersc.gov/users/software/debugging-and-profiling/craypat/). You need to access hardware counters for getting the flops; this is not a software feature. Counting flops is so last century, though, and nobody talks about this these days. I am usually just interested in getting the shortest possible walltime for a given problem I want to solve with my Uschi evolution code in the Cactus framework.
Uschi
Uschi T. Gamma Assistant to the SC, TAPIR, Caltech
*From:* Jose Fiestas Iquira jafiestas@lbl.gov *To:* Ian Hinder ian.hinder@aei.mpg.de; users@einsteintoolkit.org *Sent:* Monday, March 26, 2012 8:58 PM
*Subject:* Re: [Users] cactus performance
Hello,
Thanks for the information. I had a look at the benchmarks and profiling links on the Cactus website, and would like to find out the way to calculate Flops performed by a Cactus application. Are they some numbers published? In one of the papers you sent to me I found timings of McLachlan (which is my application), but I could not find Flops per simulation.
It seems to me I could be able to run my application using .par-files prepared for benchmarking? Please correct me if I am wrong.
My goal is to have an idea of the Flops expected by McLachlan runs on larger machines (few thousands or cores).
Thanks, Jose
2012/3/21 Frank Loeffler knarf@cct.lsu.edu
On Wed, Mar 21, 2012 at 01:15:05PM +0100, Ian Hinder wrote:
Frank, do you know what the following links have changed into?
http://www.cct.lsu.edu/xirel/ http://www.cct.lsu.edu/CCT-TR/CCT-TR-2008-5
Good question, I made an inquiry.
Frank
-----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.10 (GNU/Linux)
iQIcBAEBCAAGBQJPaelqAAoJEOkzpip+I59kRocP/RZl0R9MEiM5r/HfB+JZkwXs cA3eXKJeqiHtzsmqqruH/AI7zcR8zpq+BsrWHmtJQecqp28JqoV1+G7V8Z7cB/os 4pI0KFQQ6npRmnWytBSLrextanvLooqFgEB62S78MsyucTgzJwVX1AOkDuYQpHPO 3V6LQO7Jmcj5m2OoVr74mI/IFDermdwE/84dyjbz5tRrKPTug7qjfjFrfKT9yysK unA6oKl9zEdm1KZvDKzPTPh6hCcI1NNC+uRoNMnBu3XQprgNvZ/N8iDomrodCjhO iqKy5DVBhZVJhh//dWYuTzRu/l0MPMv4WML0OJmeLgbYerW2m+17lwAgumtccqwk OOuGTUqVfF9VuHJQ17jZ4jVyAjl9pIHENJCDZjmwNBDONinUC3Igjza3YlSHOCzU 1QUU1eomrSvA1oKThfaTKj+vpd+IjncucmQAtVW5rsiMjKQoOMajyR13kMJ6e8Y6 uooQr8wXJBJqiN48QBS3XvGTixvAeXmT+bQtBpocyr6vN4vEPBfwWymTn2sfd/XA vd/7Rps/IH9x3bbkk3eqCHiHrxvqPBfNdm/KbI20SZUT88XbFawuL7l1X3UgKCLv fjRNPFxtHQsz3CYk/z6YllKgf6xp34XYshtGnl6LE9lS2If8GOXIMR+MdgGJndre FXx7kXx6Oj4TLYsqOE0f =vnYk -----END PGP SIGNATURE-----
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
BTW, I want to run at least one Cactus iteration (I am not sure how to measure the time of one iteration), where all floating point operations are calculated. I need this number in order to have an idea of the FLOPs I obtain with my runs in another machine.
Do you know a good way to measure Flops in Cactus? Probably using software installed in LONI machines?
Thanks, Jose
On Wed, Mar 28, 2012 at 8:31 PM, Jose Fiestas Iquira jafiestas@lbl.govwrote:
Hello, Regarding performance. I am willing to run McLachlan shortly using CrayPat and I am setting the simulation time here in par/qc0-mclachlan.par, like this:
Cactus::terminate = "time" Cactus::cctk_final_time = .1
usually this time was 100.
Am I missing something else in setting the time for a short simulation? I could not find it in the documentation.
Thanks, Jose
On Mon, Mar 26, 2012 at 10:44 PM, Jose Fiestas Iquira jafiestas@lbl.govwrote:
Dear Uschi,
Thanks. I am right now trying CrayPat. I learned about it some days ago and tried my own N-cody code with CrayPat. I am using now the McLachlan Cactus thorn for my applications.
I am just not sure about the way to compile the code before running CrayPat. I was using simfactory for compilation and now I am using gmake directly. Since compilation takes some time, I am waiting for it to finish. Let see if it works.
The reason I am interested in Flops is because I want to know if my Cactus application will scale together with my own N-body runs in large machines and reach Petascale.
Btw, did you try CrayPat with Cactus? You use gmake for compiling?
Best, Jose
On Mon, Mar 26, 2012 at 10:20 PM, Ursula Gamma uschigamma@yahoo.comwrote:
Dear Jose,
I think the best is you use a performance checking tool coming with your computer. For example PAPI or CrayPat (on your lbl machine: http://www.nersc.gov/users/software/debugging-and-profiling/craypat/). You need to access hardware counters for getting the flops; this is not a software feature. Counting flops is so last century, though, and nobody talks about this these days. I am usually just interested in getting the shortest possible walltime for a given problem I want to solve with my Uschi evolution code in the Cactus framework.
Uschi
Uschi T. Gamma Assistant to the SC, TAPIR, Caltech
*From:* Jose Fiestas Iquira jafiestas@lbl.gov *To:* Ian Hinder ian.hinder@aei.mpg.de; users@einsteintoolkit.org *Sent:* Monday, March 26, 2012 8:58 PM
*Subject:* Re: [Users] cactus performance
Hello,
Thanks for the information. I had a look at the benchmarks and profiling links on the Cactus website, and would like to find out the way to calculate Flops performed by a Cactus application. Are they some numbers published? In one of the papers you sent to me I found timings of McLachlan (which is my application), but I could not find Flops per simulation.
It seems to me I could be able to run my application using .par-files prepared for benchmarking? Please correct me if I am wrong.
My goal is to have an idea of the Flops expected by McLachlan runs on larger machines (few thousands or cores).
Thanks, Jose
2012/3/21 Frank Loeffler knarf@cct.lsu.edu
On Wed, Mar 21, 2012 at 01:15:05PM +0100, Ian Hinder wrote:
Frank, do you know what the following links have changed into?
http://www.cct.lsu.edu/xirel/ http://www.cct.lsu.edu/CCT-TR/CCT-TR-2008-5
Good question, I made an inquiry.
Frank
-----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.10 (GNU/Linux)
iQIcBAEBCAAGBQJPaelqAAoJEOkzpip+I59kRocP/RZl0R9MEiM5r/HfB+JZkwXs cA3eXKJeqiHtzsmqqruH/AI7zcR8zpq+BsrWHmtJQecqp28JqoV1+G7V8Z7cB/os 4pI0KFQQ6npRmnWytBSLrextanvLooqFgEB62S78MsyucTgzJwVX1AOkDuYQpHPO 3V6LQO7Jmcj5m2OoVr74mI/IFDermdwE/84dyjbz5tRrKPTug7qjfjFrfKT9yysK unA6oKl9zEdm1KZvDKzPTPh6hCcI1NNC+uRoNMnBu3XQprgNvZ/N8iDomrodCjhO iqKy5DVBhZVJhh//dWYuTzRu/l0MPMv4WML0OJmeLgbYerW2m+17lwAgumtccqwk OOuGTUqVfF9VuHJQ17jZ4jVyAjl9pIHENJCDZjmwNBDONinUC3Igjza3YlSHOCzU 1QUU1eomrSvA1oKThfaTKj+vpd+IjncucmQAtVW5rsiMjKQoOMajyR13kMJ6e8Y6 uooQr8wXJBJqiN48QBS3XvGTixvAeXmT+bQtBpocyr6vN4vEPBfwWymTn2sfd/XA vd/7Rps/IH9x3bbkk3eqCHiHrxvqPBfNdm/KbI20SZUT88XbFawuL7l1X3UgKCLv fjRNPFxtHQsz3CYk/z6YllKgf6xp34XYshtGnl6LE9lS2If8GOXIMR+MdgGJndre FXx7kXx6Oj4TLYsqOE0f =vnYk -----END PGP SIGNATURE-----
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
users@lists.einsteintoolkit.org