Hello all,
reading the docs at http://simfactory.org/info/documentation/userguide/_auto/commands/run.html#c... it seems as if the "run" subcommand requires a --recover option to recover from a checkpoint. However looking at the simfactory generated RunScripts, this option is not used (though my runs seem to recover fine).
Since I have to do some lower-level fooling around with simfactory runs: is this option still required? Or does simfactory by now find checkpoints automatically and links them into the new folder?
In particular I am wondering if I can use a single qsub script
--8<-- #!/bins/h # PBS ... # PBS ... sim run mysim --walltime 24:00:00 --procs 42 --num-threads 13 sim cleanup mysim qsub $myself --8<--
and have it re-submit itself again and again and have simfactory properly create new output folders, link the checkpoints files from segment to segment and run with the options.
Yours, Roland
Roland
No, the --recover option should not be necessary unless you want to recover from a particular restart.
The "sim run" command is supposed to look at what actions "sim submit" did, and perform those that are missing. This includes creating the directory structure. I don't expect this to be well tested. Linking to checkpoint files cannot be done by "sim submit", so it is always "sim run" that does this.
You seem to be intending to chain jobs together. I assume you are aware that Simfactory already offers a facility for this? This would submit several the jobs simultaneously with respective dependencies. You can also pre-submit more jobs to extend this chain. Thus "sim submit" instead of "qsub" should do the same thing, and/or you can also run the "sim submit" from the command line (adding, say, ten more restarts) if needed.
The logic that triggers pre-submission is that you require more wall time than the queue limit. In this case, the job is broken into several restarts.
-erik
PS: I hope that my description above does not deviate too much from reality.
On Mon, Apr 30, 2012 at 5:46 PM, Roland Haas roland.haas@physics.gatech.edu wrote:
Hello all,
reading the docs at http://simfactory.org/info/documentation/userguide/_auto/commands/run.html#c... it seems as if the "run" subcommand requires a --recover option to recover from a checkpoint. However looking at the simfactory generated RunScripts, this option is not used (though my runs seem to recover fine).
Since I have to do some lower-level fooling around with simfactory runs: is this option still required? Or does simfactory by now find checkpoints automatically and links them into the new folder?
In particular I am wondering if I can use a single qsub script
--8<-- #!/bins/h # PBS ... # PBS ... sim run mysim --walltime 24:00:00 --procs 42 --num-threads 13 sim cleanup mysim qsub $myself --8<--
and have it re-submit itself again and again and have simfactory properly create new output folders, link the checkpoints files from segment to segment and run with the options.
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Hello Erik,
No, the --recover option should not be necessary unless you want to recover from a particular restart.
Thank you. That is good to know.
The "sim run" command is supposed to look at what actions "sim submit" did, and perform those that are missing. This includes creating the directory structure. I don't expect this to be well tested. Linking to checkpoint files cannot be done by "sim submit", so it is always "sim run" that does this.
ok.
You seem to be intending to chain jobs together. I assume you are aware that Simfactory already offers a facility for this? This would submit several the jobs simultaneously with respective dependencies. You can also pre-submit more jobs to extend this chain. Thus "sim submit" instead of "qsub" should do the same thing, and/or you can also run the "sim submit" from the command line (adding, say, ten more restarts) if needed.
I am using presubmission for other jobs (though it has been fragile occasionally in not finding checkpoints and not warning about it, but I am not quite sure how I triggered it. It's one of those "magic" things that almost always work it seems unless one does something unexpected.).
I cannot use the presubmission facility straightforwardly (I believe) since I currently lump a number of independent simulations into a single queue item (to get a bit of higher run priority on kraken due to the higher core count).
The simplest way right now seems to use "sim submit ... --mdbkey=submit '/bin/bash @SCRIPTFILE@ &'" in my qsub script to start all sub-simulations which means I can simply re-use the very same qsub script again at the end after waiting for all simulations to finish.
The logic that triggers pre-submission is that you require more wall time than the queue limit. In this case, the job is broken into several restarts.
It also happens if I submit a job to to a simulation that is already in queue or running, does it not? Or would I explicitly have to add a --presubmit (or similar) flag to the submit sub-command?
Yours, Roland
PS: I hope that my description above does not deviate too much from reality.
Well I can always try and see how reality compares to the design specifications. Any difference is a bug in the specifications...
On Tue, May 1, 2012 at 11:57 AM, Roland Haas roland.haas@physics.gatech.edu wrote:
I cannot use the presubmission facility straightforwardly (I believe) since I currently lump a number of independent simulations into a single queue item (to get a bit of higher run priority on kraken due to the higher core count).
There should be an explicit facility for this. I'm not sure whether this should be part of Simfactory or not. This would also be useful for running many benchmarks in parallel etc.
The logic that triggers pre-submission is that you require more wall time than the queue limit. In this case, the job is broken into several restarts.
It also happens if I submit a job to to a simulation that is already in queue or running, does it not? Or would I explicitly have to add a --presubmit (or similar) flag to the submit sub-command?
It should work as-is, without an explicit presubmit flag.
-erik
users@lists.einsteintoolkit.org