Periodically when continuing a run, I get the error
Error: job id is negative
Aborting Simfactory
It is fixed by deleting the symlink to the previous active directory and then resubmitting.
Although that fix works, when I want to submit a job multiple times at once (so that when one output is finished running, the next one starts automatically), this error prevents me from doing that.
Anyone have any ideas on what might be the underlying cause of this issue and how to actually fix it?
Thanks!
Deborah
-----------------------------------------
Deborah Ferguson
Graduate Student
Georgia Institute of Technology
Center for Relativistic Astrophysics
Boggs 1-66
Hello Deborah,
the issue is that the run command did not set a jobid entry in the properties.ini file.
This needs to be fixed in simfactory in this file: lib/simrestart.py which needs to be modified such that the userRun command stores a jobid in the same way that the submit command does in line 765.
This would be a bit of a hack since self.userRun is analogous to self.userSubmit and self.userSubmit does not store the jobid (sself.ubmit does, but run cannot store the self.jobid since it exeutes from *within* the job). To make it worse, self.userRun does not even know the jobid since the jobid is returned by the submit command, the submit entry in the machine.ini file, which is normally qsub/sbatch/whatnot but configured to by something else that just starts the SubmitScript as a background process and returns the process ID for machines without schedulers.
So something like:
'jobid': os.getpid()
in line 945 (just after 'pbsSimulationName': pbsSimulationName) probably does the trick. It's a hack though since it makes assumptions what the job id will be (namely my process ID) and is dangerous in case process IDs are re-used (there's only 65k of them so this does happen).
Yours, Roland
Periodically when continuing a run, I get the error
Error: job id is negative
Aborting Simfactory
It is fixed by deleting the symlink to the previous active directory and then resubmitting.
Although that fix works, when I want to submit a job multiple times at once (so that when one output is finished running, the next one starts automatically), this error prevents me from doing that.
Anyone have any ideas on what might be the underlying cause of this issue and how to actually fix it?
Thanks!
Deborah
Deborah Ferguson
Graduate Student
Georgia Institute of Technology
Center for Relativistic Astrophysics
Boggs 1-66
Dear all,
I'm a new user, trying to get familiar with Llama
I wanted to run the TOV example parameter file with a multipatch grid, so I changed it. Unfortunately it does not work out and I can't see where the error is coming from. The original static_tov.par worked fine on my machine.
Does anyone see what I messed up?
Thanks a lot!
Best regards,
Severin
Hello Severin,
in your err file, it says:
--8<-- /home/severin/simulations/static_tov_llama/SIMFACTORY/exe/cactus_sim: error while loading shared libraries: libgsl.so.19: cannot open shared object file: No such file or directory --8<--
which indicates that the gsl library is not found.
Why this is the case can have multiple different reasons.
If you are doing this on a cluster where you used a "module" command to load a GSL module then most likely you have to use the same "module" command in your runscript before the mpirun line.
It is also possible the the libraries present on the login node (where you compiled) do not match the ones present on the compute nodes.
You may want to take your .err file output and show it either to a colleague at Tuebingen who has used the cluster before or, failing that, show it one of the cluster's support persons who may be able to advise on how to proceed.
On a laptop this typically does not happen since there the linker checks whether the library is present at link time and (unless eg a system update changes the libraries in between) the same library will be found at runtime. Though for some Linux distributions (eg RedHat based ones) one may also have to use a module command. Finally if your gls library is installed in a "strange" place you may have to add that location to your LD_LIBRARY_PATH variable (at least as a temporary fix).
Yours, Roland
Dear all,
I'm a new user, trying to get familiar with Llama
I wanted to run the TOV example parameter file with a multipatch grid, so I changed it. Unfortunately it does not work out and I can't see where the error is coming from. The original static_tov.par worked fine on my machine.
Does anyone see what I messed up?
Thanks a lot!
Best regards,
Severin
users@lists.einsteintoolkit.org