Christian Ott, Peter Diener, and I tracked down a set of checkpointing/recovery inconsistencies in the Einstein Toolkit. This means that many variables had different values after checkpointing and then recovering them, which should not be the case. We were careful not to use MPI, OpenMP, or any kind of compiler optimisations, so these inconsistencies represent real changes.
We found a series of problems, and we have now one possible correction. Unfortunately, this requires changes to several thorns, mostly to schedule.ccl declarations, but also to Carpet. In addition, we find that it is impossible to re-calculate PseudoEvolution variables after recovery -- they have to be checkpointed. In other words, variables such as ADMBase, TmunuBase, or other variables with 3 timelevels need to be checkpointed to be consistent.
(Many of the inconsistencies are small, and are caused by differences in the discretisation error. That is, they will vanish in the continuum limit. However, one of the design principles of the Einstein Toolkit is that results should be as much as possible independent of the number of MPI processes, OpenMP threads, and checkpointing/recovery. Disabling checkpointing of such variables could be implemented depending on a parameter, which would also need to ensure that these variables are then recalculated -- with slightly different values -- after recovery. In particular, this would require a new schedule bin in Carpet that is executed after recovery for all timelevels, or alternatively a new schedule option that requests this behaviour from Carpet.)
The necessary changes to correct these inconsistencies are in detail:
1. All variables with 3 timelevels need to be checkpointed. This affects ADMBase, TmunuBase, ML_ADMConstraints, ML_ADMQuantities, and the BSSN constraints in ML_BSSN. Since this depends on the number of timelevels, care has to be taken to make the right decision at run time.
2. The schedule group MoL_PseudoEvolution must not be scheduled in post_recover_variables -- since these variables are checkpointed, they do not need to be recalculated.
3. Carpet cannot execute the post_recover_variables bin on the past timelevels, since applying boundary conditions to past timelevels includes prolongation, and the time interpolation necessary for this would have a different discretisation error.
4. Carpet traversed the PostRestrict bin in the incorrect order. It traversed from finest to coarsest (the same order in which restriction has to be applied), but because fine grid boundaries may be interpolate from coarse grid boundaries, this bin must be traversed from coarsest to finest.
5. The scheduling of GRHydro and HydroBase is off in certain places.
6. The scheduling of NaNChecker leaves the NaNmask uninitialised after recovery.
7. The scheduling of various Kranc generated thorns is off in certain places.
8. Certain schedule items requiring the ADMBase variables are executed in the incorrect order, i.e. before the ADMBase variables were available.
I attach a diff of a possible set of corrections. However: - This diff does not correct handling the TmunuBase variables -- they are still not checkpointed if they have multiple timelevels. I believe this should be done by thorn TmunuBase itself. - Similarly, handling checkpointing of the ADMBase variables should be moved from ML_BSSN_Helper to ADMBase itself. - I think a new schedule group MoL_PseudoEvolutionBoundaries would make sense, as this would simplify the schedule of thorns which use MoL_PseudoEvolution.
-erik
This is really good news! Thank you and congratulations. It was really a long-standing problem (and hard to investigate).
Questions:
- In the diff, which version of carpet do the changes refer to? (All?)
- When do you plan to commit the changes to the repositories?
Thanks,
Luca
On 20/7/11 11:21 AM, Erik Schnetter wrote:
Christian Ott, Peter Diener, and I tracked down a set of checkpointing/recovery inconsistencies in the Einstein Toolkit. This means that many variables had different values after checkpointing and then recovering them, which should not be the case. We were careful not to use MPI, OpenMP, or any kind of compiler optimisations, so these inconsistencies represent real changes.
We found a series of problems, and we have now one possible correction. Unfortunately, this requires changes to several thorns, mostly to schedule.ccl declarations, but also to Carpet. In addition, we find that it is impossible to re-calculate PseudoEvolution variables after recovery -- they have to be checkpointed. In other words, variables such as ADMBase, TmunuBase, or other variables with 3 timelevels need to be checkpointed to be consistent.
(Many of the inconsistencies are small, and are caused by differences in the discretisation error. That is, they will vanish in the continuum limit. However, one of the design principles of the Einstein Toolkit is that results should be as much as possible independent of the number of MPI processes, OpenMP threads, and checkpointing/recovery. Disabling checkpointing of such variables could be implemented depending on a parameter, which would also need to ensure that these variables are then recalculated -- with slightly different values -- after recovery. In particular, this would require a new schedule bin in Carpet that is executed after recovery for all timelevels, or alternatively a new schedule option that requests this behaviour from Carpet.)
The necessary changes to correct these inconsistencies are in detail:
- All variables with 3 timelevels need to be checkpointed. This
affects ADMBase, TmunuBase, ML_ADMConstraints, ML_ADMQuantities, and the BSSN constraints in ML_BSSN. Since this depends on the number of timelevels, care has to be taken to make the right decision at run time.
- The schedule group MoL_PseudoEvolution must not be scheduled in
post_recover_variables -- since these variables are checkpointed, they do not need to be recalculated.
- Carpet cannot execute the post_recover_variables bin on the past
timelevels, since applying boundary conditions to past timelevels includes prolongation, and the time interpolation necessary for this would have a different discretisation error.
- Carpet traversed the PostRestrict bin in the incorrect order. It
traversed from finest to coarsest (the same order in which restriction has to be applied), but because fine grid boundaries may be interpolate from coarse grid boundaries, this bin must be traversed from coarsest to finest.
The scheduling of GRHydro and HydroBase is off in certain places.
The scheduling of NaNChecker leaves the NaNmask uninitialised after recovery.
The scheduling of various Kranc generated thorns is off in certain places.
Certain schedule items requiring the ADMBase variables are executed
in the incorrect order, i.e. before the ADMBase variables were available.
I attach a diff of a possible set of corrections. However:
- This diff does not correct handling the TmunuBase variables -- they
are still not checkpointed if they have multiple timelevels. I believe this should be done by thorn TmunuBase itself.
- Similarly, handling checkpointing of the ADMBase variables should be
moved from ML_BSSN_Helper to ADMBase itself.
- I think a new schedule group MoL_PseudoEvolutionBoundaries would
make sense, as this would simplify the schedule of thorns which use MoL_PseudoEvolution.
-erik
On Tue, Jul 19, 2011 at 10:43 PM, Luca Baiotti baiotti@ile.osaka-u.ac.jp wrote:
This is really good news! Thank you and congratulations. It was really a long-standing problem (and hard to investigate).
Questions:
- In the diff, which version of carpet do the changes refer to? (All?)
This refers to the Mercurial version.
- When do you plan to commit the changes to the repositories?
Very soon.
-erik
Thanks,
Luca
On 20/7/11 11:21 AM, Erik Schnetter wrote:
Christian Ott, Peter Diener, and I tracked down a set of checkpointing/recovery inconsistencies in the Einstein Toolkit. This means that many variables had different values after checkpointing and then recovering them, which should not be the case. We were careful not to use MPI, OpenMP, or any kind of compiler optimisations, so these inconsistencies represent real changes.
We found a series of problems, and we have now one possible correction. Unfortunately, this requires changes to several thorns, mostly to schedule.ccl declarations, but also to Carpet. In addition, we find that it is impossible to re-calculate PseudoEvolution variables after recovery -- they have to be checkpointed. In other words, variables such as ADMBase, TmunuBase, or other variables with 3 timelevels need to be checkpointed to be consistent.
(Many of the inconsistencies are small, and are caused by differences in the discretisation error. That is, they will vanish in the continuum limit. However, one of the design principles of the Einstein Toolkit is that results should be as much as possible independent of the number of MPI processes, OpenMP threads, and checkpointing/recovery. Disabling checkpointing of such variables could be implemented depending on a parameter, which would also need to ensure that these variables are then recalculated -- with slightly different values -- after recovery. In particular, this would require a new schedule bin in Carpet that is executed after recovery for all timelevels, or alternatively a new schedule option that requests this behaviour from Carpet.)
The necessary changes to correct these inconsistencies are in detail:
- All variables with 3 timelevels need to be checkpointed. This
affects ADMBase, TmunuBase, ML_ADMConstraints, ML_ADMQuantities, and the BSSN constraints in ML_BSSN. Since this depends on the number of timelevels, care has to be taken to make the right decision at run time.
- The schedule group MoL_PseudoEvolution must not be scheduled in
post_recover_variables -- since these variables are checkpointed, they do not need to be recalculated.
- Carpet cannot execute the post_recover_variables bin on the past
timelevels, since applying boundary conditions to past timelevels includes prolongation, and the time interpolation necessary for this would have a different discretisation error.
- Carpet traversed the PostRestrict bin in the incorrect order. It
traversed from finest to coarsest (the same order in which restriction has to be applied), but because fine grid boundaries may be interpolate from coarse grid boundaries, this bin must be traversed from coarsest to finest.
The scheduling of GRHydro and HydroBase is off in certain places.
The scheduling of NaNChecker leaves the NaNmask uninitialised after recovery.
The scheduling of various Kranc generated thorns is off in certain places.
Certain schedule items requiring the ADMBase variables are executed
in the incorrect order, i.e. before the ADMBase variables were available.
I attach a diff of a possible set of corrections. However:
- This diff does not correct handling the TmunuBase variables -- they
are still not checkpointed if they have multiple timelevels. I believe this should be done by thorn TmunuBase itself.
- Similarly, handling checkpointing of the ADMBase variables should be
moved from ML_BSSN_Helper to ADMBase itself.
- I think a new schedule group MoL_PseudoEvolutionBoundaries would
make sense, as this would simplify the schedule of thorns which use MoL_PseudoEvolution.
-erik
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 20 Jul 2011, at 03:21, Erik Schnetter wrote:
Christian Ott, Peter Diener, and I tracked down a set of checkpointing/recovery inconsistencies in the Einstein Toolkit. This means that many variables had different values after checkpointing and then recovering them, which should not be the case. We were careful not to use MPI, OpenMP, or any kind of compiler optimisations, so these inconsistencies represent real changes.
We found a series of problems, and we have now one possible correction. Unfortunately, this requires changes to several thorns, mostly to schedule.ccl declarations, but also to Carpet. In addition, we find that it is impossible to re-calculate PseudoEvolution variables after recovery -- they have to be checkpointed. In other words, variables such as ADMBase, TmunuBase, or other variables with 3 timelevels need to be checkpointed to be consistent.
(Many of the inconsistencies are small, and are caused by differences in the discretisation error. That is, they will vanish in the continuum limit. However, one of the design principles of the Einstein Toolkit is that results should be as much as possible independent of the number of MPI processes, OpenMP threads, and checkpointing/recovery. Disabling checkpointing of such variables could be implemented depending on a parameter, which would also need to ensure that these variables are then recalculated -- with slightly different values -- after recovery. In particular, this would require a new schedule bin in Carpet that is executed after recovery for all timelevels, or alternatively a new schedule option that requests this behaviour from Carpet.)
Did you find cases where the evolved variables had different values after recovery, or only analysis variables? Can you estimate the order of the errors introduced? Are the errors introduced only at refinement boundaries?
- The scheduling of various Kranc generated thorns is off in certain places.
Could you elaborate on this? Does Kranc itself have to be modified?
- I think a new schedule group MoL_PseudoEvolutionBoundaries would
make sense, as this would simplify the schedule of thorns which use MoL_PseudoEvolution.
The problem with this is that if you have two functions scheduled in MoL_PseudoEvolution and the second one uses values computed in the first, these values will not be correct in the boundaries.
On Wed, Jul 20, 2011 at 8:28 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 20 Jul 2011, at 03:21, Erik Schnetter wrote:
Christian Ott, Peter Diener, and I tracked down a set of checkpointing/recovery inconsistencies in the Einstein Toolkit. This means that many variables had different values after checkpointing and then recovering them, which should not be the case. We were careful not to use MPI, OpenMP, or any kind of compiler optimisations, so these inconsistencies represent real changes.
We found a series of problems, and we have now one possible correction. Unfortunately, this requires changes to several thorns, mostly to schedule.ccl declarations, but also to Carpet. In addition, we find that it is impossible to re-calculate PseudoEvolution variables after recovery -- they have to be checkpointed. In other words, variables such as ADMBase, TmunuBase, or other variables with 3 timelevels need to be checkpointed to be consistent.
(Many of the inconsistencies are small, and are caused by differences in the discretisation error. That is, they will vanish in the continuum limit. However, one of the design principles of the Einstein Toolkit is that results should be as much as possible independent of the number of MPI processes, OpenMP threads, and checkpointing/recovery. Disabling checkpointing of such variables could be implemented depending on a parameter, which would also need to ensure that these variables are then recalculated -- with slightly different values -- after recovery. In particular, this would require a new schedule bin in Carpet that is executed after recovery for all timelevels, or alternatively a new schedule option that requests this behaviour from Carpet.)
Did you find cases where the evolved variables had different values after recovery, or only analysis variables? Can you estimate the order of the errors introduced? Are the errors introduced only at refinement boundaries?
Yes. 1e-4. Yes.
- The scheduling of various Kranc generated thorns is off in certain places.
Could you elaborate on this? Does Kranc itself have to be modified?
I'm not sure yet. It probably doesn't have to be modified, but I'd like to introduce PseudoEvolutionBoundaries, and would like to use this group in Kranc (for the *_bc_group).
- I think a new schedule group MoL_PseudoEvolutionBoundaries would
make sense, as this would simplify the schedule of thorns which use MoL_PseudoEvolution.
The problem with this is that if you have two functions scheduled in MoL_PseudoEvolution and the second one uses values computed in the first, these values will not be correct in the boundaries.
Boundary conditions would be scheduled in both groups. There are some times when only boundary conditions are needed, and this is when PseudoEvolutionBoundaries comes in. The alternative is to schedule the boundary conditions explicitly in postrestrictinitial, postrestrict, postregridinitial, and postregrid, which is cumbersome.
-erik
After a bit of clean-up, this is the final patch I intend to commit. In addition to what was discussed earlier, this also introduces a new MoL schedule group MoL_PostStepModify:
Phyics thorns can apply enforce constraints in PostStepModify. The difference between PostStep and PostStepModify is that PostStep is scheduled at many other occasions, whereas the PostStepModify is only scheduled during evolution.
Please comment.
-erik
Hello Erik, all,
After a bit of clean-up, this is the final patch I intend to commit. In addition to what was discussed earlier, this also introduces a new MoL schedule group MoL_PostStepModify:
Phyics thorns can apply enforce constraints in PostStepModify. The difference between PostStep and PostStepModify is that PostStep is scheduled at many other occasions, whereas the PostStepModify is only scheduled during evolution.
So PostStepModify should contain eg. evolution boundary conditions and dissipation (eg.) while PostStep would contain symmetry boundary conditions, con2prim etc? Would that also mean that routines scheduled in PostStep are supposed to be idempotent (ie. applying them twice does not change the values on the grid)? Note that the later excludes the current implementation of Con2Prim (I think) since they will do at least grhydro_count_min Newton iterations (and therefore always change values).
Otherwise I like the idea of separating enforcement of constraints from evolution-type calculations.
I have a comment for this part of the patch:
--8<-- Index: arrangements/EinsteinEvolve/GRHydro/schedule.ccl =================================================================== --- arrangements/EinsteinEvolve/GRHydro/schedule.ccl (revision 251) +++ arrangements/EinsteinEvolve/GRHydro/schedule.ccl (working copy) @@ -983,9 +983,9 @@ } "Reset the atmosphere" }
-schedule group HydroBase_Boundaries IN MoL_Evolution AFTER MoL_Step -{ -} "HydroBase Boundary conditions group" +#schedule group HydroBase_Boundaries IN MoL_Evolution AFTER MoL_Step +#{ +#} "HydroBase Boundary conditions group" --8<--
the reason why this extra Boundary group was scheduled was that ResetAtmoshphere runs in MoL_Evolution AFTER MoL_Step. ResetAtmoshpere uses values in space_mask to reset some points, however space_mask was only set for the interior points (since it depends in part on the RHS) so in order to have consistent data one has to SYNC and apply boundary conditions (though con2prim is not required I think). Is this still true?
Yours, Roland
On Mon, Aug 1, 2011 at 10:37 AM, Roland Haas roland.haas@physics.gatech.edu wrote:
Hello Erik, all,
After a bit of clean-up, this is the final patch I intend to commit. In addition to what was discussed earlier, this also introduces a new MoL schedule group MoL_PostStepModify:
Phyics thorns can apply enforce constraints in PostStepModify. The difference between PostStep and PostStepModify is that PostStep is scheduled at many other occasions, whereas the PostStepModify is only scheduled during evolution.
So PostStepModify should contain eg. evolution boundary conditions and dissipation (eg.) while PostStep would contain symmetry boundary conditions, con2prim etc? Would that also mean that routines scheduled in PostStep are supposed to be idempotent (ie. applying them twice does not change the values on the grid)? Note that the later excludes the current implementation of Con2Prim (I think) since they will do at least grhydro_count_min Newton iterations (and therefore always change values).
Almost: MoL_PostStepModify would contain those evolution steps that are independent of MoL. Dissipation is applied to the RHS, so it doesn't go into this group. Currently, only enforcing the BSSN constraints is scheduled here.
Yes, MoL_PostStep routines are supposed to be idempotent, since they are applied "randomly" at various occations, e.g. after regridding. If there is a regridding step that doesn't modify level L, then MoL_PostStep may still be executed on level L, and if a routine is not idempotent, the results will differ. I was thinking of con2prim myself. In a production run, with MPI and OpenMP and compiler optimisations, things will differ anyway "randomly", but it may be nice to have a certain operating mode for con2prim which makes it idempotent, maybe only for a simple EOS.
Otherwise I like the idea of separating enforcement of constraints from evolution-type calculations.
I have a comment for this part of the patch:
--8<-- Index: arrangements/EinsteinEvolve/GRHydro/schedule.ccl =================================================================== --- arrangements/EinsteinEvolve/GRHydro/schedule.ccl (revision 251) +++ arrangements/EinsteinEvolve/GRHydro/schedule.ccl (working copy) @@ -983,9 +983,9 @@ } "Reset the atmosphere" }
-schedule group HydroBase_Boundaries IN MoL_Evolution AFTER MoL_Step -{ -} "HydroBase Boundary conditions group" +#schedule group HydroBase_Boundaries IN MoL_Evolution AFTER MoL_Step +#{ +#} "HydroBase Boundary conditions group" --8<--
the reason why this extra Boundary group was scheduled was that ResetAtmoshphere runs in MoL_Evolution AFTER MoL_Step. ResetAtmoshpere uses values in space_mask to reset some points, however space_mask was only set for the interior points (since it depends in part on the RHS) so in order to have consistent data one has to SYNC and apply boundary conditions (though con2prim is not required I think). Is this still true?
Thanks. I added HydroBase_Boundaries at this location again. Note that this group may sync and hence may be expensive -- skipping this group unless necessary may be a good idea.
-erik
users@lists.einsteintoolkit.org