Present: Lucas (Minutes), Steven R Brandt, Johnny Tsao, Maxwell Rizzo, Roland Hass (Chair), Erik Schnetter.
# Vote to change the cadence of Einstein Toolkit releases. * After intense discussions it was decided (by the voting of those who were present) that the releases shall now be **every 12 months, every May** * It was agreed that this decision may be revisited in the future, in case the community finds it necessary.
# Discussions on minor releases and continuous integration * The change of cadence prompted intense discussions on whether we should do point releases with important bug fixes in between releases. Several points were proposed, in particular the following ideas were proposed 1. Do not perform point releases, as this would be equivalent to having two releases a year. 2. Leave the decision to do point releases to the current release manager. 3. Adopt a CI (continuous integration) scheme that would allow us to know if any of the changes on the master branch are breaking in the machines that we support. 4. In order to ameliorate difficulties of running tests (even if automated) on certain machines, it was proposed that we adopt a system of champions and tiers. Each machine would have a champion, responsible for making sure that the ETK works correctly on that system. These would be Tier 1 machines, with full support and working guarantees. Other machines, where champions were out of touch or unable to get the toolkit working in time, would be called Tier 2 machines. * Erik suggested that we postpone this discussion for the next meetings, as the topic is broad and requires more thought. All members agreed.
# Reproducibility issues with BNS * Wolfgang Tichy has obtained different results with different MPI ranks on BNS runs and reported them to Roland. Roland also did a bit of testing with the BNS gallery example, writing checkpoints at iteration 0 and 1. He sees differences between runs with 4 and 8 MPI ranks. Zach reminded Roland that Newton's method exhibits chaotic behavior, which his group observed on Illinois GMHD and may explain the results. Erik countered the argument saying we should assume that a bug is more likely. Roland says he needs to perform more tests, but he suspects that boundary condition interactions may be causing the problems. Zach suggested that Roland determines if the hydro quantities change before the metric changes, as hydro changes would take some time to propagate to metric components. Erik suggested checking different time levels. Erik also suggests comparing RHSs. Finally, Zach recommends trying running it through Valgrind, to check accidental memory issues. Roland took note of all suggestions and will implement them.
# Upcoming Einstein Toolkit Release * NewRadX: Even though the code is completed, it still needs reviews and patches to be approved. * Z4C: Pending review by reviewers. * Cauchy characteristic extraction for Spectre CCE: Deborah Ferguson, who is the champion of this code, says that work on it is on her agenda but she is currently busy finishing a paper for her group. * Cosmology codes: Roland has reached out to Hayley Macpherson to see if she has any code that she would like to include. She has not reached back yet. * Gallery examples for more modules in the ETK: Roland reports that there are two such examples in the works based on Canuda for evolving the Einstein + various fields systems. Roland reports summer student projects for new gallery examples may not happen as the summer student budget was cut.
# Unanswered question on the mailing list
No questions.
# Open tickets sorted by update time
2845: Small change for fixing C++ namespaces. 2837: Help request with not much progress since last week.
# Tickets ready for review * Roland requests for reviewers, particularly those interested in CarpetX, to help on the review of the first few tickets such as 2845, which are relatively small and easy to review.
Hello all,
# Reproducibility issues with BNS
- Wolfgang Tichy has obtained different results with different MPI ranks on
Sorry, my bad. Correction: Michal Pirog (who is working with Wolfgang) noticed these.
Yours, Roland
Regarding the BNS issues, is it something due to when you recover from checkpoint?
In Whisky we had an issue related to that due to Con2Prim running after recovery from checkpoint.
E.g., you run a simulation from iteration 0 to iteration 24 (made up numbers just to explain myself). You try the same simulation, but this time you run from 0 to 12, checkpoint, recover from checkpoint and run from 12 to 24. The results between these two simulations at iteration 24 were slightly different
The issue was in the Newton Raphson being called by Con2Prim and using a different initial guess when recovering from checkpoint. During a run it uses the previous values of the primitive variables as an initial guess, but these are not available when restarting from checkpoint. The solution, as far as I remember, was not to run Con2Prim during recovery from checkpoint (since the primitive variables are anyway recovered from the checkpoint files).
Cheers, Bruno
Il giorno gio 16 gen 2025 alle ore 19:59 Roland Haas rhaas@illinois.edu ha scritto:
Hello all,
# Reproducibility issues with BNS
- Wolfgang Tichy has obtained different results with different MPI
ranks on Sorry, my bad. Correction: Michal Pirog (who is working with Wolfgang) noticed these.
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://pgp.mit.edu . _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Hello all, Bruno,
Before this takes on a live of its own:
I re-did my own tests and found no more differences after a single Euler step.
The issue that caused me to think there were differences before were:
* I had looked at the wrong column in norms output and mistaken an "average" norm for a "maximum" norm (and the average will change due to different orders of values being summed up)
* I still think I see a difference in values for points that are at the same location on different refinement levels (which should also not be in a vertex centered run), but those seem the same no matter whether I use 4 or 8 MPI ranks (also a bug if true, but of different kind)
Regarding the BNS issues, is it something due to when you recover from checkpoint?
Maybe, I don't know.
In Whisky we had an issue related to that due to Con2Prim running after recovery from checkpoint.
Yes, I had pointed it out in the past to Michal.
For those one can set:
MoL::run_MoL_PostStep_in_Post_Recover_Variables = "no"
which as Bruno said, disables the extra con2prim after checkpoint recovery.
Yours, Roland
users@lists.einsteintoolkit.org