#2888: race condition writing properties.ini in simfactory
Reporter: Roland Haas
Status: new
Milestone:
Version:
Type: bug
Priority: major
Component: SimFactory
Comment (by Roland Haas):
Note that not writing `properties.ini` is not a sufficient solution since even without writing there remains the issue that a job could start before `jobid` has been recorded and thus the `@JOB_ID@` replacement is invalid.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2888/race-condition-wr…
#2888: race condition writing properties.ini in simfactory
Reporter: Roland Haas
Status: new
Milestone:
Version:
Type: bug
Priority: major
Component: SimFactory
Currently bot the `submit()` as well as the `run()`function in `lib/simrestart.py` write \(update\) the file `output-NNNN/properties.ini`. In particular `submit()` does so _after_ submitting to record the `jobid` field while `run()` does so before running to record the `checkpointing` value.
This means that if \(eg on SLURM where jobs start quickly, an particular when the `submit` command contains a `sleep 5` to slow down job submission\) there can be race condition:
1. `submit` writes initial copy of `properties.ini` lacking `jobid`
2. `submit` submits the job to SLURM
3. job starts and run reads `properties.ini`
4. `submit` writes updated `properties.ini` with `jobid`
5. `run` writes `properties.ini` with `checkpointing`
at this point `jobid` is lost from `properties.ini`
Fixes would be:
* add a lock for `properties.ini` which `submit` only release once it has done its final update
* submit jobs in “held” state \(`sbatch --hold`\) and only release them once the final update of properties.ini has happened
both will require updates to simfactory’s code. The first option would work without updates to the machine ini files, the second will require extra entries to tell simfactory how to release a job. This could also be used to implement `hold` and `release` commands in simfactory. The first option will require fallback code in case the lock file is left behind \(e.g. wait no longer than 1 minute before assuming a stale lock and going ahead anyway\). The lock file will require at least some level for POSIX conformance from the file system \(namely that file creation and removal is an atomic operation cluster-wide\).
This most likely is the reason for occasional strange failures on clusters with jobid being unset.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2888/race-condition-wr…
#2482: Inconsistency in Simfactory submission scripts
Reporter: Steven R. Brandt
Status: new
Milestone:
Version:
Type: bug
Priority: major
Component: SimFactory
Comment (by Roland Haas):
I think this can be closed. Each cluster is too different and usually maintained by different people to really hope for consistent files between clusters.
@{557058:1671c5c3-29cc-4e83-9850-a152d33a6235} ?
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/2482/inconsistency-in-…
#127: Restarts should have hard link to executable
Reporter: Erik Schnetter
Status: open
Milestone:
Version:
Type: bug
Priority: minor
Component: SimFactory
Comment (by Roland Haas):
Hard links will face the same issues as the executable cache concerning file systems that do not allow hard links across directories \(BeeGFS maybe also some Lustre setups that want to use multiple metadata servers, not sure\).
Having a fixed copy of the executable that was used is useful when one updates the executable during a run. Currently this is a manual intervention.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/127/restarts-should-ha…
#1721: Simfactory's test mechanism should use softlinks instead of rsync
Reporter: Erik Schnetter
Status: open
Milestone:
Version: development version
Type: enhancement
Priority: minor
Component: SimFactory
Comment (by Roland Haas):
As pointed out in [https://bitbucket.org/einsteintoolkit/tickets/issues/1721/simfactorys-test-… not all clusters allow access to the file system that contains the source code from a running job.
So copies are required there.
To avoid the complication of a cache or determining if symbolic links would work, I suggest we keep the copying.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/1721/simfactorys-test-…
#1321: Collisions in executable cache directory
Reporter: Ian Hinder
Status: new
Milestone:
Version: development version
Type: bug
Priority: minor
Component: SimFactory
Comment (by Roland Haas):
Caching the executable \(currently about 1.1 GB when compiling all thorns and with debug symbols\) in order to reduce disk space usage becomes mostly moot once the first checkpoint is written: even the lowest resolution checkpoint for a vacuum run of mine is 2.0 GB per checkpoint file, and there are 64 checkpoint files per checkpoint.
So it really only matters for things like test runs \(or the testsuites\) where no checkpoints are being written.
The cache also uses hard-links across directories which not all file systems like / support \(BeeGFS is one of them\).
Do we need the cache? Removing it would remove all issues associated with it.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/1321/collisions-in-exe…
#1327: Delay subsequent restarts in the case of certain problems
Reporter: Ian Hinder
Status: duplicate
Milestone:
Version:
Type: enhancement
Priority: minor
Component: SimFactory
Changes (by Roland Haas):
status: duplicate (was new)
Comment (by Roland Haas):
Duplicate of #1286.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/1327/delay-subsequent-…
#1327: Delay subsequent restarts in the case of certain problems
Reporter: Ian Hinder
Status: new
Milestone:
Version:
Type: enhancement
Priority: minor
Component: SimFactory
Comment (by Roland Haas):
Mostly a duplicate of #1286
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/1327/delay-subsequent-…
#1335: SimFactory should abort if there are no checkpoint files when submitting an existing configuration
Reporter: Ian Hinder
Status: wontfix
Milestone:
Version:
Type: enhancement
Priority: minor
Component: SimFactory
Changes (by Roland Haas):
status: wontfix (was new)
When submitting a simulation which already contains at least one restart, simfactory should abort if there are no checkpoint files available. This likely means that something went wrong. Starting the simulation again is always the wrong thing to do in this case, as it will waste CPU time and might go unnoticed.
**Keyword:**
Comment (by Roland Haas):
As Ian stated, there can realistically be no checkpoint files even when one wants to "continue" a simulation.
--
Ticket URL: https://bitbucket.org/einsteintoolkit/tickets/issues/1335/simfactory-should…