Hi. Can someone explain to me why ML_BSSN calls SelectBoundConds within MoL_PostStep? It seems like the kind of once-off routine that would happen near the start of a simulation, rather than something that has to be performed every single timestep.
I wouldn't mind, but while trying to understand why ML_BSSN was evolving so slowly on one of our machines, I looked at the TimerReport files, and saw that SelectBoundConds was taking *much* more time (like 20 times as long) than the actual RHS calculation routines.
Bernard
Hello Bernard,
Hi. Can someone explain to me why ML_BSSN calls SelectBoundConds within MoL_PostStep? It seems like the kind of once-off routine that would happen near the start of a simulation, rather than something that has to be performed every single timestep.
In the "new" boundary/symmetry interface (ie using thorn Boundary and Symbase) one has to Select the variables for boundaries each time before ApplyBCs is scheduled. There is a routine in Boundaries that clears the selection.
I wouldn't mind, but while trying to understand why ML_BSSN was evolving so slowly on one of our machines, I looked at the TimerReport files, and saw that SelectBoundConds was taking *much* more time (like 20 times as long) than the actual RHS calculation routines.
The long time is most likely caused by the fact that the boundary selection routine tends to be the one calling SYNC which means it is the one that does an MPI wait (if there is load imbalance) and communicates data for buffer zone prolongation etc.
Yours, Roland
Hi Roland, and thanks for your reply. I'm still a bit confused, I confess (see below) ...
On 2/16/13 12:38 AM, "Roland Haas" roland.haas@physics.gatech.edu wrote:
Hello Bernard,
Hi. Can someone explain to me why ML_BSSN calls SelectBoundConds within MoL_PostStep? It seems like the kind of once-off routine that would happen near the start of a simulation, rather than something that has to be performed every single timestep.
In the "new" boundary/symmetry interface (ie using thorn Boundary and Symbase) one has to Select the variables for boundaries each time before ApplyBCs is scheduled. There is a routine in Boundaries that clears the selection.
OK; but can you tell me *why* this is? Why should we ever have to re-specify what kind of boundary conditions we use during a simulation, any more than we re-specify the evolution equations? Perhaps I don't really understand what "select the variables" means here.
I wouldn't mind, but while trying to understand why ML_BSSN was evolving so slowly on one of our machines, I looked at the TimerReport files, and saw that SelectBoundConds was taking *much* more time (like 20 times as long) than the actual RHS calculation routines.
The long time is most likely caused by the fact that the boundary selection routine tends to be the one calling SYNC which means it is the one that does an MPI wait (if there is load imbalance) and communicates data for buffer zone prolongation etc.
So it might be spending most of the time waiting for other cores to catch up? But if it's really waiting for prior routines to finish on other processors, then on the handful of cores where SBC appears significantly *quicker* than usual (e.g. ~50,000 seconds instead of ~100,000) I should see earlier routines taking correspondingly *longer*, right? But I don't.
I'm attaching TimerReport files for two cores on the same (128-core) evolution. Core 000 is typical. Line 3591 (the most up-to-date instance of "large" SBC behaviour) shows about 100K seconds spent cumulatively over the simulation so far. Core 052 shows only about half as much time used in the same routine, but I can't see what other EVOL routines might be taking up the slack.
(Note, BTW, that what I'm running isn't vanilla ML_BSSN, but a locally modified version called MH_BSSN. The scheduling and most routines are almost identical to McLachlan)
Bernard
Yours, Roland
Bernard
Did you set the parameters that Ian Hinder suggested? This will make interpreting the timing output much easier.
Time is not only spent in routines, but also in infrastructure tasks not listed here (regridding etc.). Ian Hinder implemented a tree-form timer output that makes it much easier to see how much time is spent on what task. I don't recall the options to select this -- Ian, is this Carpet::output_timer_tree_every and friends?
-erik
On Mon, Feb 18, 2013 at 2:59 PM, Kelly, Bernard J. (GSFC-660.0)[UNIVERSITY OF MARYLAND BALTIMORE COUNTY] bernard.j.kelly@nasa.gov wrote:
Hi Roland, and thanks for your reply. I'm still a bit confused, I confess (see below) ...
On 2/16/13 12:38 AM, "Roland Haas" roland.haas@physics.gatech.edu wrote:
Hello Bernard,
Hi. Can someone explain to me why ML_BSSN calls SelectBoundConds within MoL_PostStep? It seems like the kind of once-off routine that would happen near the start of a simulation, rather than something that has to be performed every single timestep.
In the "new" boundary/symmetry interface (ie using thorn Boundary and Symbase) one has to Select the variables for boundaries each time before ApplyBCs is scheduled. There is a routine in Boundaries that clears the selection.
OK; but can you tell me *why* this is? Why should we ever have to re-specify what kind of boundary conditions we use during a simulation, any more than we re-specify the evolution equations? Perhaps I don't really understand what "select the variables" means here.
I wouldn't mind, but while trying to understand why ML_BSSN was evolving so slowly on one of our machines, I looked at the TimerReport files, and saw that SelectBoundConds was taking *much* more time (like 20 times as long) than the actual RHS calculation routines.
The long time is most likely caused by the fact that the boundary selection routine tends to be the one calling SYNC which means it is the one that does an MPI wait (if there is load imbalance) and communicates data for buffer zone prolongation etc.
So it might be spending most of the time waiting for other cores to catch up? But if it's really waiting for prior routines to finish on other processors, then on the handful of cores where SBC appears significantly *quicker* than usual (e.g. ~50,000 seconds instead of ~100,000) I should see earlier routines taking correspondingly *longer*, right? But I don't.
I'm attaching TimerReport files for two cores on the same (128-core) evolution. Core 000 is typical. Line 3591 (the most up-to-date instance of "large" SBC behaviour) shows about 100K seconds spent cumulatively over the simulation so far. Core 052 shows only about half as much time used in the same routine, but I can't see what other EVOL routines might be taking up the slack.
(Note, BTW, that what I'm running isn't vanilla ML_BSSN, but a locally modified version called MH_BSSN. The scheduling and most routines are almost identical to McLachlan)
Bernard
Yours, Roland
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 19 Feb 2013, at 19:19, Erik Schnetter schnetter@cct.lsu.edu wrote:
Bernard
Did you set the parameters that Ian Hinder suggested? This will make interpreting the timing output much easier.
Time is not only spent in routines, but also in infrastructure tasks not listed here (regridding etc.). Ian Hinder implemented a tree-form timer output that makes it much easier to see how much time is spent on what task. I don't recall the options to select this -- Ian, is this Carpet::output_timer_tree_every and friends?
Yes. That one displays the timer tree of the current process to standard output, and is useful for a quick overview. I recommend enabling it in all simulations. Unfortunately, it doesn't do a reduction across processes. You can also output the top N timers, not in tree form, using n_top_timers; these *are* reduced across processes, but because many of the timers overlap (e.g. some timers contain other timers), it is harder to get a good overview. You can also output an XML file per process containing the full hierarchical timer tree, which you can process in some other tool to do the reduction across processes (I have code in Mathematica to do this, if you are interested).
Ideally, I would like the stdout timer tree to include reductions across processes (min/max/average), and to have a version of the XML output which was also similarly reduced. The reason this was not easy to do was that you could have different timers existing on each process, so some reductions wouldn't work. It might be that this is no longer the case, and that implementing the reductions would be more straightforward now. Erik?
-erik
On Mon, Feb 18, 2013 at 2:59 PM, Kelly, Bernard J. (GSFC-660.0)[UNIVERSITY OF MARYLAND BALTIMORE COUNTY] bernard.j.kelly@nasa.gov wrote: Hi Roland, and thanks for your reply. I'm still a bit confused, I confess (see below) ...
On 2/16/13 12:38 AM, "Roland Haas" roland.haas@physics.gatech.edu wrote:
Hello Bernard,
Hi. Can someone explain to me why ML_BSSN calls SelectBoundConds within MoL_PostStep? It seems like the kind of once-off routine that would happen near the start of a simulation, rather than something that has to be performed every single timestep.
In the "new" boundary/symmetry interface (ie using thorn Boundary and Symbase) one has to Select the variables for boundaries each time before ApplyBCs is scheduled. There is a routine in Boundaries that clears the selection.
OK; but can you tell me *why* this is? Why should we ever have to re-specify what kind of boundary conditions we use during a simulation, any more than we re-specify the evolution equations? Perhaps I don't really understand what "select the variables" means here.
I wouldn't mind, but while trying to understand why ML_BSSN was evolving so slowly on one of our machines, I looked at the TimerReport files, and saw that SelectBoundConds was taking *much* more time (like 20 times as long) than the actual RHS calculation routines.
The long time is most likely caused by the fact that the boundary selection routine tends to be the one calling SYNC which means it is the one that does an MPI wait (if there is load imbalance) and communicates data for buffer zone prolongation etc.
So it might be spending most of the time waiting for other cores to catch up? But if it's really waiting for prior routines to finish on other processors, then on the handful of cores where SBC appears significantly *quicker* than usual (e.g. ~50,000 seconds instead of ~100,000) I should see earlier routines taking correspondingly *longer*, right? But I don't.
I'm attaching TimerReport files for two cores on the same (128-core) evolution. Core 000 is typical. Line 3591 (the most up-to-date instance of "large" SBC behaviour) shows about 100K seconds spent cumulatively over the simulation so far. Core 052 shows only about half as much time used in the same routine, but I can't see what other EVOL routines might be taking up the slack.
(Note, BTW, that what I'm running isn't vanilla ML_BSSN, but a locally modified version called MH_BSSN. The scheduling and most routines are almost identical to McLachlan)
Bernard
Yours, Roland
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/ _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On Mon, Feb 18, 2013 at 01:59:41PM -0600, Kelly, Bernard J. (GSFC-660.0)[UNIVERSITY OF MARYLAND BALTIMORE COUNTY] wrote:
evolution. Core 000 is typical. Line 3591 (the most up-to-date instance of "large" SBC behaviour) shows about 100K seconds spent cumulatively over the simulation so far. Core 052 shows only about half as much time used in the same routine, but I can't see what other EVOL routines might be taking up the slack.
Some thoughts:
Could it be that their position is different, e.g., is one of them a corner and the other someplace in the middle of the domain? These two cases would have to handle a different amount of communication. However, I would naively assume that core000 would get a corner and thus, should be faster, but that doesn't seem to be the case. Maybe you use a symmetry and one of them just happens to be at such a boundary and the other doesn't? Also, the actual number of ghost points does influence the result and could be different for different decompositions. Carpet's decomposition tries to distribute the number of evolved points equally first while keeping the boxes cube-like only as second criteria. I wouldn't expect communication to be distributed equally.
Also, sometimes the network connection in clusters can differ quite a bit depending on the nodes you get and their connection. Not all nodes might have the same connectivity to all other nodes. However, I didn't find this to be as bad as you describe either.
Frank
[re-sent, with smaller attachment]
Hi Roland, and thanks for your reply. I'm still a bit confused, I confess (see below) ...
On 2/16/13 12:38 AM, "Roland Haas" roland.haas@physics.gatech.edu wrote:
Hello Bernard,
Hi. Can someone explain to me why ML_BSSN calls SelectBoundConds within MoL_PostStep? It seems like the kind of once-off routine that would happen near the start of a simulation, rather than something that has to be performed every single timestep.
In the "new" boundary/symmetry interface (ie using thorn Boundary and Symbase) one has to Select the variables for boundaries each time before ApplyBCs is scheduled. There is a routine in Boundaries that clears the selection.
OK; but can you tell me *why* this is? Why should we ever have to re-specify what kind of boundary conditions we use during a simulation, any more than we re-specify the evolution equations? Perhaps I don't really understand what "select the variables" means here.
I wouldn't mind, but while trying to understand why ML_BSSN was evolving so slowly on one of our machines, I looked at the TimerReport files, and saw that SelectBoundConds was taking *much* more time (like 20 times as long) than the actual RHS calculation routines.
The long time is most likely caused by the fact that the boundary selection routine tends to be the one calling SYNC which means it is the one that does an MPI wait (if there is load imbalance) and communicates data for buffer zone prolongation etc.
So it might be spending most of the time waiting for other cores to catch up? But if it's really waiting for prior routines to finish on other processors, then on the handful of cores where SBC appears significantly *quicker* than usual (e.g. ~50,000 seconds instead of ~100,000) I should see earlier routines taking correspondingly *longer*, right? But I don't.
I'm attaching TimerReport files for two cores on the same (128-core) evolution. Core 000 is typical. Line 184 (the most up-to-date instance of "large" SBC behaviour) shows about 100K seconds spent cumulatively over the simulation so far. Core 052 shows only about half as much time used in the same routine, but I can't see what other EVOL routines might be taking up the slack.
(Note, BTW, that what I'm running isn't vanilla ML_BSSN, but a locally modified version called MH_BSSN. The scheduling and most routines are almost identical to McLachlan)
Bernard
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
On 18 Feb 2013, at 21:11, "Kelly, Bernard J. (GSFC-660.0)[UNIVERSITY OF MARYLAND BALTIMORE COUNTY]" bernard.j.kelly@nasa.gov wrote:
[re-sent, with smaller attachment]
Hi Roland, and thanks for your reply. I'm still a bit confused, I confess (see below) ...
On 2/16/13 12:38 AM, "Roland Haas" roland.haas@physics.gatech.edu wrote:
Hello Bernard,
Hi. Can someone explain to me why ML_BSSN calls SelectBoundConds within MoL_PostStep? It seems like the kind of once-off routine that would happen near the start of a simulation, rather than something that has to be performed every single timestep.
In the "new" boundary/symmetry interface (ie using thorn Boundary and Symbase) one has to Select the variables for boundaries each time before ApplyBCs is scheduled. There is a routine in Boundaries that clears the selection.
OK; but can you tell me *why* this is? Why should we ever have to re-specify what kind of boundary conditions we use during a simulation, any more than we re-specify the evolution equations? Perhaps I don't really understand what "select the variables" means here.
Take a look at the documentation for thorn Boundary (http://einsteintoolkit.org/documentation/ThornDoc/CactusBase/Boundary/docume...). The "selection" process is part of the API of the thorn. Maybe one of the old-timers can explain why it was done this way.
I wouldn't mind, but while trying to understand why ML_BSSN was evolving so slowly on one of our machines, I looked at the TimerReport files, and saw that SelectBoundConds was taking *much* more time (like 20 times as long) than the actual RHS calculation routines.
The long time is most likely caused by the fact that the boundary selection routine tends to be the one calling SYNC which means it is the one that does an MPI wait (if there is load imbalance) and communicates data for buffer zone prolongation etc.
So it might be spending most of the time waiting for other cores to catch up?
If you look at timer output just for one process, you will almost certainly reach erroneous conclusions due to things like this. I recommend to look at the output on all processes (yes, performance profiling is hard).
But if it's really waiting for prior routines to finish on other processors, then on the handful of cores where SBC appears significantly *quicker* than usual (e.g. ~50,000 seconds instead of ~100,000) I should see earlier routines taking correspondingly *longer*, right? But I don't.
It may also be that timings change significantly from one iteration to the next. Have you set your CPU affinity settings correctly?
I recommend to set the parameters
Carpet::schedule_barriers = yes Carpet::sync_barriers = yes
This will insert an MPI barrier before and after each scheduled function call and sync. Then you can rely on the timings of the individual functions, and also see how much time is spent waiting to catch up (i.e. in load imbalance). At the moment, the function timers for functions which do communication will include time spent waiting for the other process to catch up.
I'm attaching TimerReport files for two cores on the same (128-core) evolution. Core 000 is typical. Line 184 (the most up-to-date instance of "large" SBC behaviour) shows about 100K seconds spent cumulatively over the simulation so far. Core 052 shows only about half as much time used in the same routine, but I can't see what other EVOL routines might be taking up the slack.
(Note, BTW, that what I'm running isn't vanilla ML_BSSN, but a locally modified version called MH_BSSN. The scheduling and most routines are almost identical to McLachlan)
Bernard
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
<TimerReports_LATEST_BJK.tgz>_______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Hi Ian (and Frank and Erik). Thanks for the further insight on the profiling.
[Please ignore the new mail that just came through with the 400KB attachment. That was my first attempt that was held for moderation because of the attachment size. Then I sent the slimmed-down attachments, but this was still in the pipeline.]
I was looking at *all* the processor outputs (that is, all the TimerReport_XXXXXX files), but not necessarily at all fields in all of them. I concentrated on the CCTK_EVOL section of the report, and then only looked closely at discrepancies between a sample "longer SelectBoundCond" processor and each of the five or six "shorter SelectBoundcond" processors. I suppose to do a more complete job, I'd have to start scripting ...
Anyway, I *hadn't* been using those profiling parameters before, so my conclusions were probably dodgy as you say. After your reply I re-enabled them and restarted the run. Since it's so slow, I'm now looking at the TimerReports from earlier in the new run, and no longer see any discrepancies between different processors (that is, there don't seem to be any "shorter SelectBoundcond" processors any more).
So if *all* the processors are showing essentially the same information, and the "schedule_barriers" and "sync_barriers" are in place, then there's no significant load imbalance? And yet it is slow as hell ...
I'm now testing with the actual repository McLachlan instead.
Bernard
On 2/18/13 3:40 PM, "Ian Hinder" ian.hinder@aei.mpg.de wrote:
On 18 Feb 2013, at 21:11, "Kelly, Bernard J. (GSFC-660.0)[UNIVERSITY OF MARYLAND BALTIMORE COUNTY]" bernard.j.kelly@nasa.gov wrote:
[re-sent, with smaller attachment]
Hi Roland, and thanks for your reply. I'm still a bit confused, I confess (see below) ...
I wouldn't mind, but while trying to understand why ML_BSSN was evolving so slowly on one of our machines, I looked at the TimerReport files, and saw that SelectBoundConds was taking *much* more time (like 20 times as long) than the actual RHS calculation routines.
The long time is most likely caused by the fact that the boundary selection routine tends to be the one calling SYNC which means it is the one that does an MPI wait (if there is load imbalance) and communicates data for buffer zone prolongation etc.
So it might be spending most of the time waiting for other cores to catch up?
If you look at timer output just for one process, you will almost certainly reach erroneous conclusions due to things like this. I recommend to look at the output on all processes (yes, performance profiling is hard).
But if it's really waiting for prior routines to finish on other processors, then on the handful of cores where SBC appears significantly *quicker* than usual (e.g. ~50,000 seconds instead of ~100,000) I should see earlier routines taking correspondingly *longer*, right? But I don't.
It may also be that timings change significantly from one iteration to the next. Have you set your CPU affinity settings correctly?
I recommend to set the parameters
Carpet::schedule_barriers = yes Carpet::sync_barriers = yes
This will insert an MPI barrier before and after each scheduled function call and sync. Then you can rely on the timings of the individual functions, and also see how much time is spent waiting to catch up (i.e. in load imbalance). At the moment, the function timers for functions which do communication will include time spent waiting for the other process to catch up.
I'm attaching TimerReport files for two cores on the same (128-core) evolution. Core 000 is typical. Line 184 (the most up-to-date instance of "large" SBC behaviour) shows about 100K seconds spent cumulatively over the simulation so far. Core 052 shows only about half as much time used in the same routine, but I can't see what other EVOL routines might be taking up the slack.
(Note, BTW, that what I'm running isn't vanilla ML_BSSN, but a locally modified version called MH_BSSN. The scheduling and most routines are almost identical to McLachlan)
Bernard
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
<TimerReports_LATEST_BJK.tgz>____________________________________________ ___ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
On Tue, Feb 19, 2013 at 1:24 PM, Kelly, Bernard J. (GSFC-660.0)[UNIVERSITY OF MARYLAND BALTIMORE COUNTY] bernard.j.kelly@nasa.gov wrote:
Hi Ian (and Frank and Erik). Thanks for the further insight on the profiling.
[Please ignore the new mail that just came through with the 400KB attachment. That was my first attempt that was held for moderation because of the attachment size. Then I sent the slimmed-down attachments, but this was still in the pipeline.]
I was looking at *all* the processor outputs (that is, all the TimerReport_XXXXXX files), but not necessarily at all fields in all of them. I concentrated on the CCTK_EVOL section of the report, and then only looked closely at discrepancies between a sample "longer SelectBoundCond" processor and each of the five or six "shorter SelectBoundcond" processors. I suppose to do a more complete job, I'd have to start scripting ...
Anyway, I *hadn't* been using those profiling parameters before, so my conclusions were probably dodgy as you say. After your reply I re-enabled them and restarted the run. Since it's so slow, I'm now looking at the TimerReports from earlier in the new run, and no longer see any discrepancies between different processors (that is, there don't seem to be any "shorter SelectBoundcond" processors any more).
So if *all* the processors are showing essentially the same information, and the "schedule_barriers" and "sync_barriers" are in place, then there's no significant load imbalance? And yet it is slow as hell ...
With schedule barriers, load imbalance is hidden in these barriers. That is, you would need to measure how much time each process spends in these barriers. I expect that some processes will spend 0s there, while others will spend 50,000s there. That would be your load imbalance.
-erik
I'm now testing with the actual repository McLachlan instead.
Bernard
On 2/18/13 3:40 PM, "Ian Hinder" ian.hinder@aei.mpg.de wrote:
On 18 Feb 2013, at 21:11, "Kelly, Bernard J. (GSFC-660.0)[UNIVERSITY OF MARYLAND BALTIMORE COUNTY]" bernard.j.kelly@nasa.gov wrote:
[re-sent, with smaller attachment]
Hi Roland, and thanks for your reply. I'm still a bit confused, I confess (see below) ...
I wouldn't mind, but while trying to understand why ML_BSSN was evolving so slowly on one of our machines, I looked at the TimerReport files, and saw that SelectBoundConds was taking *much* more time (like 20 times as long) than the actual RHS calculation routines.
The long time is most likely caused by the fact that the boundary selection routine tends to be the one calling SYNC which means it is the one that does an MPI wait (if there is load imbalance) and communicates data for buffer zone prolongation etc.
So it might be spending most of the time waiting for other cores to catch up?
If you look at timer output just for one process, you will almost certainly reach erroneous conclusions due to things like this. I recommend to look at the output on all processes (yes, performance profiling is hard).
But if it's really waiting for prior routines to finish on other processors, then on the handful of cores where SBC appears significantly *quicker* than usual (e.g. ~50,000 seconds instead of ~100,000) I should see earlier routines taking correspondingly *longer*, right? But I don't.
It may also be that timings change significantly from one iteration to the next. Have you set your CPU affinity settings correctly?
I recommend to set the parameters
Carpet::schedule_barriers = yes Carpet::sync_barriers = yes
This will insert an MPI barrier before and after each scheduled function call and sync. Then you can rely on the timings of the individual functions, and also see how much time is spent waiting to catch up (i.e. in load imbalance). At the moment, the function timers for functions which do communication will include time spent waiting for the other process to catch up.
I'm attaching TimerReport files for two cores on the same (128-core) evolution. Core 000 is typical. Line 184 (the most up-to-date instance of "large" SBC behaviour) shows about 100K seconds spent cumulatively over the simulation so far. Core 052 shows only about half as much time used in the same routine, but I can't see what other EVOL routines might be taking up the slack.
(Note, BTW, that what I'm running isn't vanilla ML_BSSN, but a locally modified version called MH_BSSN. The scheduling and most routines are almost identical to McLachlan)
Bernard
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
<TimerReports_LATEST_BJK.tgz>____________________________________________ ___ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 19 Feb 2013, at 20:16, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Tue, Feb 19, 2013 at 1:24 PM, Kelly, Bernard J. (GSFC-660.0)[UNIVERSITY OF MARYLAND BALTIMORE COUNTY] bernard.j.kelly@nasa.gov wrote: Hi Ian (and Frank and Erik). Thanks for the further insight on the profiling.
[Please ignore the new mail that just came through with the 400KB attachment. That was my first attempt that was held for moderation because of the attachment size. Then I sent the slimmed-down attachments, but this was still in the pipeline.]
I was looking at *all* the processor outputs (that is, all the TimerReport_XXXXXX files), but not necessarily at all fields in all of them. I concentrated on the CCTK_EVOL section of the report, and then only looked closely at discrepancies between a sample "longer SelectBoundCond" processor and each of the five or six "shorter SelectBoundcond" processors. I suppose to do a more complete job, I'd have to start scripting ...
Anyway, I *hadn't* been using those profiling parameters before, so my conclusions were probably dodgy as you say. After your reply I re-enabled them and restarted the run. Since it's so slow, I'm now looking at the TimerReports from earlier in the new run, and no longer see any discrepancies between different processors (that is, there don't seem to be any "shorter SelectBoundcond" processors any more).
So if *all* the processors are showing essentially the same information, and the "schedule_barriers" and "sync_barriers" are in place, then there's no significant load imbalance? And yet it is slow as hell ...
With schedule barriers, load imbalance is hidden in these barriers. That is, you would need to measure how much time each process spends in these barriers. I expect that some processes will spend 0s there, while others will spend 50,000s there. That would be your load imbalance.
When I added the sync barriers, I added timers on all the barriers. You should see timer entries named ".../Barrier". Do you see these, and are they taking a lot of time? The timer names are hierarchical, so you should be able to see which function barriers are causing the slowdown.
When I have done tests using schedule barriers, they did not impose a huge penalty like the one you are describing. Maybe 30%, no more.
Hi Ian.
The TimerReport.XXXXXX.txt files have no instance of "Barrier" (any capitalisation) appearing. Does it depend on anything else apart from the parameters you've specified? I'm attaching one process's TimerReport (only the last time output, for compactness), and the associated parfile here.
Bernard
From: Ian Hinder <ian.hinder@aei.mpg.demailto:ian.hinder@aei.mpg.de> Date: Tuesday, February 19, 2013 4:16 PM To: Erik Schnetter <schnetter@cct.lsu.edumailto:schnetter@cct.lsu.edu> Cc: Bernard Kelly <bernard.j.kelly@nasa.govmailto:bernard.j.kelly@nasa.gov>, "users@einsteintoolkit.orgmailto:users@einsteintoolkit.org" <users@einsteintoolkit.orgmailto:users@einsteintoolkit.org> Subject: Re: [Users] logic of scheduling SelectBoundConds in McLachlan?
On 19 Feb 2013, at 20:16, Erik Schnetter <schnetter@cct.lsu.edumailto:schnetter@cct.lsu.edu> wrote:
On Tue, Feb 19, 2013 at 1:24 PM, Kelly, Bernard J. (GSFC-660.0)[UNIVERSITY OF MARYLAND BALTIMORE COUNTY] <bernard.j.kelly@nasa.govmailto:bernard.j.kelly@nasa.gov> wrote: Hi Ian (and Frank and Erik). Thanks for the further insight on the profiling.
[Please ignore the new mail that just came through with the 400KB attachment. That was my first attempt that was held for moderation because of the attachment size. Then I sent the slimmed-down attachments, but this was still in the pipeline.]
I was looking at *all* the processor outputs (that is, all the TimerReport_XXXXXX files), but not necessarily at all fields in all of them. I concentrated on the CCTK_EVOL section of the report, and then only looked closely at discrepancies between a sample "longer SelectBoundCond" processor and each of the five or six "shorter SelectBoundcond" processors. I suppose to do a more complete job, I'd have to start scripting ...
Anyway, I *hadn't* been using those profiling parameters before, so my conclusions were probably dodgy as you say. After your reply I re-enabled them and restarted the run. Since it's so slow, I'm now looking at the TimerReports from earlier in the new run, and no longer see any discrepancies between different processors (that is, there don't seem to be any "shorter SelectBoundcond" processors any more).
So if *all* the processors are showing essentially the same information, and the "schedule_barriers" and "sync_barriers" are in place, then there's no significant load imbalance? And yet it is slow as hell ...
With schedule barriers, load imbalance is hidden in these barriers. That is, you would need to measure how much time each process spends in these barriers. I expect that some processes will spend 0s there, while others will spend 50,000s there. That would be your load imbalance.
When I added the sync barriers, I added timers on all the barriers. You should see timer entries named ".../Barrier". Do you see these, and are they taking a lot of time? The timer names are hierarchical, so you should be able to see which function barriers are causing the slowdown.
When I have done tests using schedule barriers, they did not impose a huge penalty like the one you are describing. Maybe 30%, no more.
-- Ian Hinder http://numrel.aei.mpg.de/people/hinder
On Mon, Feb 18, 2013 at 02:11:29PM -0600, Kelly, Bernard J. (GSFC-660.0)[UNIVERSITY OF MARYLAND BALTIMORE COUNTY] wrote:
OK; but can you tell me *why* this is?
I am not sure either, but the point is less that you respecify them (because that is fast), but that you have to sync before applying them - and these functions just happen to be the place this is done.
Frank
On Mon, Feb 18, 2013 at 3:11 PM, Kelly, Bernard J. (GSFC-660.0)[UNIVERSITY OF MARYLAND BALTIMORE COUNTY] bernard.j.kelly@nasa.gov wrote:
[re-sent, with smaller attachment]
Hi Roland, and thanks for your reply. I'm still a bit confused, I confess (see below) ...
On 2/16/13 12:38 AM, "Roland Haas" roland.haas@physics.gatech.edu wrote:
Hello Bernard,
Hi. Can someone explain to me why ML_BSSN calls SelectBoundConds within MoL_PostStep? It seems like the kind of once-off routine that would happen near the start of a simulation, rather than something that has to be performed every single timestep.
In the "new" boundary/symmetry interface (ie using thorn Boundary and Symbase) one has to Select the variables for boundaries each time before ApplyBCs is scheduled. There is a routine in Boundaries that clears the selection.
OK; but can you tell me *why* this is? Why should we ever have to re-specify what kind of boundary conditions we use during a simulation, any more than we re-specify the evolution equations? Perhaps I don't really understand what "select the variables" means here.
Boundary conditions can be applied at different points during a simulation, and one may want to apply different boundary conditions at different times.
The canonical example is the conformal factor: While calculating initial conditions (e.g. with an elliptic solver), one wants to apply Robin boundary conditions; during evolution, one wants a radiative boundary condition. Similar arguments would hold if one were to enforce constraints during evolution.
-erik
users@lists.einsteintoolkit.org