Hi All,
I have some very long BH simulations to run and I'd like to checkpoint for these. I haven't really done checkpointing before. But what I know is that chekpointing information can be specified in the parameter file (for use by Cactus), and also that Simfactory does seem to have some stuff to do with or handle checkointing ( "restart-id", etc...). Of course, scheduling systems (e.g. PBSPro) at HPCs would have support for checkpointing but I don't want to use that. Probably it is only best to use that to set the walltime.
So, my main question is: Assuming I set a maximum walltime of 12 hours, and I set my simulation to dump checkpoints every 3hrs (in walltime units), how do I *restart* my job at the end of the 12 hrs using Simfactory in a way that the simulation starts off from the last checkpoint it droppped before terminating? What extra command line options should I pass to the sumbit command of SImfactory?
Below is a segment of my parfile where checkpointing information is given, provides as a sample or as a basis for anyone who would want to advise me on how such information should be given.
##### Checkpointing ######### CarpetIOHDF5::checkpoint = yes IO::checkpoint_ID = yes IO::recover = autoprobe IO::checkpoint_every = 1024 IO::out_proc_every = 2 IO::checkpoint_keep = 3 IO::checkpoint_dir = $parfile Carpet::regrid_during_recovery = no CarpetIOHDF5::use_grid_structure_from_checkpoint = yes CarpetIOHDF5::open_one_input_file_at_a_time = yes
Your advice and assistance will be highly appreciated.
Best, Dumsani
On Wed, Aug 31, 2016 at 11:02:16PM +0200, dumsani wrote:
Below is a segment of my parfile where checkpointing information is given, provides as a sample or as a basis for anyone who would want to advise me on how such information should be given.
Maybe it helps to explain what these options do.
##### Checkpointing ######### CarpetIOHDF5::checkpoint = yes
This enables checkpointing.
IO::checkpoint_ID = yes
This specifies that initial data (ID) should be checkpointed as well. So, you get a checkpoint right after initial data generation and before the first evolution step. This makes sense if ID generation takes long, or if you are interested in these data itself.
IO::recover = autoprobe
This instructs Cactus to restart from the latest checkpoint it finds, if it finds any. If it doesn't find any, it starts from initial data, like without checkpointing.
IO::checkpoint_every = 1024
Checkpoint every so many iterations. Personally, I wouldn't use this, but a setting that depends on (wall) time - but there is nothing wrong with it.
IO::out_proc_every = 2
Not specifically related to checkpointing.
IO::checkpoint_keep = 3
Keep the last 3 checkpoints, delete older versions.
IO::checkpoint_dir = $parfile
Put checkpoint files into a directory that is 'X' if the parameter file was called X.par (remove the .par extension).
Ideally, this should be accompanied by:
IO::recover_dir = $parfile
Otherwise, using the same parfile, Cactus wouldn't find the generated checkpoint files.
Carpet::regrid_during_recovery = no CarpetIOHDF5::use_grid_structure_from_checkpoint = yes
Don't change the grid structure during recovery.
CarpetIOHDF5::open_one_input_file_at_a_time = yes
Use less memory while reading files, at the possible expense of time.
Something interesting as well:
IO::checkpoint_on_terminate = "yes"
Dump a checkpoint on terminate. This enables to terminate the simulation at a certain point (e.g., just before wall-time runs out), and continue exactly where you stopped.
Frank
On 31 Aug 2016, at 23:02, dumsani g14n8326@campus.ru.ac.za wrote:
Hi All,
I have some very long BH simulations to run and I'd like to checkpoint for these. I haven't really done checkpointing before. But what I know is that chekpointing information can be specified in the parameter file (for use by Cactus), and also that Simfactory does seem to have some stuff to do with or handle checkointing ( "restart-id", etc...). Of course, scheduling systems (e.g. PBSPro) at HPCs would have support for checkpointing but I don't want to use that. Probably it is only best to use that to set the walltime.
So, my main question is: Assuming I set a maximum walltime of 12 hours, and I set my simulation to dump checkpoints every 3hrs (in walltime units), how do I *restart* my job at the end of the 12 hrs using Simfactory in a way that the simulation starts off from the last checkpoint it droppped before terminating? What extra command line options should I pass to the sumbit command of SImfactory?
You don't need anything extra; just "sim submit <simulationname>".
– If the job has completed already, the next job will be queued. – If the job is queued or running, the next job will be queued with a dependency to only start when the previous one finishes. (The dependency logic is in the submit script of the machine; it's possible that the machine you are using does not have this defined. Look for references to "chain" in the other submit scripts in case you need to add this to your own machine.)
You can also use
TerminationTrigger::max_walltime = @WALLTIME_HOURS@ TerminationTrigger::on_remaining_walltime = 30 # minutes TerminationTrigger::output_remtime_every_minutes = 30
This will cause Cactus to cleanly terminate 30 minutes before the end of the job's walltime (as a margin). If you additionally use
IO::checkpoint_on_terminate = yes
then you will get a checkpoint written. Without this, your job will be unceremoniously killed by the scheduler, leaving you with up to 3 hours of wasted computer time, possible corrupted output files, and duplicate data.
It is also convenient to use
TerminationTrigger::termination_from_file = yes TerminationTrigger::termination_file = "terminate.txt" TerminationTrigger::create_termination_file = yes
This will create a file called "terminate.txt" in the output directory. If you add a "1" to this file, Cactus will terminate immediately (and checkpoint, if you have set checkpoint_on_terminate as above). You can then resubmit the simulation if you like. This allows you to easily stop and start simulations without losing any runtime.
users@lists.einsteintoolkit.org