Hi, community,
I was doing some benchmark tests for my version of GRHydro with TimerReport and I found that a lot of time in the code is spent in SYNCs for the staggered variables. Analogous to IllinoisGRMHD I have the following lines in the schedule.ccl
{
SYNC:GRHydro::em_Ax,GRHydro::em_Ay,GRHydro::em_Az,GRHydro::em_psi6phi LANG: C } "Schedule symmetries -- Actually just a placeholder function to ensure prolongations / processor syncs are done BEFORE outer boundaries are updated."
According to the timer, this takes 3 times more than other SYNCs. Is there a reason for this? Is it related to the prolongation operators for staggered variables?
Thanks for your help!
*Dr. Luciano Combi* Postdoctoral Researcher Perimeter Institute for Theoretical Physics CITA National Fellow (U. of Guelph) ---
Hello Luciano,
Given that factor of 3 that you notice, it could be that this is mostly some latency when doing network communication.
If your memory usage allows it you can experiment with the option
CarpetLib::combine_sends = "yes"
see
https://bitbucket.org/eschnett/carpet/src/master/CarpetLib/param.ccl#lines-2...
and see if that makes things go faster.
This will send all data to all ranks at once instead of going a bit slower so may help.
Yours, Roland
Hi, community,
I was doing some benchmark tests for my version of GRHydro with TimerReport and I found that a lot of time in the code is spent in SYNCs for the staggered variables. Analogous to IllinoisGRMHD I have the following lines in the schedule.ccl
{
SYNC:GRHydro::em_Ax,GRHydro::em_Ay,GRHydro::em_Az,GRHydro::em_psi6phi LANG: C } "Schedule symmetries -- Actually just a placeholder function to ensure prolongations / processor syncs are done BEFORE outer boundaries are updated."
According to the timer, this takes 3 times more than other SYNCs. Is there a reason for this? Is it related to the prolongation operators for staggered variables?
Thanks for your help!
*Dr. Luciano Combi* Postdoctoral Researcher Perimeter Institute for Theoretical Physics CITA National Fellow (U. of Guelph)
Thanks for the tip Roland.
I have tried it and it seems I'm getting the same efficiency nevertheless.
I was wondering whether, perhaps, the prolongations of the staggered variables are somehow more expensive/inefficient?
Thanks a lot.
Luciano
On Mon, Apr 1, 2024 at 12:29 PM Roland Haas rhaas@illinois.edu wrote:
Hello Luciano,
Given that factor of 3 that you notice, it could be that this is mostly some latency when doing network communication.
If your memory usage allows it you can experiment with the option
CarpetLib::combine_sends = "yes"
see
https://bitbucket.org/eschnett/carpet/src/master/CarpetLib/param.ccl#lines-2...
and see if that makes things go faster.
This will send all data to all ranks at once instead of going a bit slower so may help.
Yours, Roland
Hi, community,
I was doing some benchmark tests for my version of GRHydro with
TimerReport
and I found that a lot of time in the code is spent in SYNCs for the staggered variables. Analogous to IllinoisGRMHD I have the following
lines
in the schedule.ccl
{
SYNC:GRHydro::em_Ax,GRHydro::em_Ay,GRHydro::em_Az,GRHydro::em_psi6phi LANG: C } "Schedule symmetries -- Actually just a placeholder function to ensure prolongations / processor syncs are done BEFORE outer
boundaries are
updated."
According to the timer, this takes 3 times more than other SYNCs. Is
there
a reason for this? Is it related to the prolongation operators for staggered variables?
Thanks for your help!
*Dr. Luciano Combi* Postdoctoral Researcher Perimeter Institute for Theoretical Physics CITA National Fellow (U. of Guelph)
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://pgp.mit.edu .
Hello Luciano,
Thanks for the tip Roland.
I have tried it and it seems I'm getting the same efficiency nevertheless.
Yeah, it was a bit of a long shot.
I was wondering whether, perhaps, the prolongations of the staggered variables are somehow more expensive/inefficient?
I could not think of anything that would make it more expensive. It is just a slightly different interpolation stencil, but has the same size overall.
Could in principle be less efficient if the functions are less well optimized.
My guess would normally be that its related to the communication time and not so much to the actual computation time of the interpolation.
Maybe Zach Etienne, whose IllinoisGRMHD code initially required this sort of interpolation, recalls any issues?
If not you could try and run with all verbose options on
Carpet::verbose = yes Carpet::veryverbose = yes
CarpetLib::verbose = yes CarpetLib::commstate_verbose = yes
and possibly
Carpet::sync_barriers = yes
to get more consistent (but possibly slower than normal) timing results.
Note that these will produce a lot of log output so you only want to run a very few timesteps.
@Erik: just to confirm, if one has a:
SYNC: Ax, Ay
then those two synchronizations happen in parallel and not first Ax then a MPI_WaitAll then Ay then a MPI_WaitAll (I had thought they'd be one after the other, but looking at CarpetLib's code in commstate and gdata it seems they all happen at once)?
Yours, Roland
users@lists.einsteintoolkit.org