#141: Don't output cleanup errors for perl-style simulations
------------------------+---------------------------------------------------
Reporter: eschnett | Owner: mthomas
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
While cleaning, up, don't output lots of errors if perl-style simulations
are encountered. Instead, skip them silently. (Find a good way to
distinguish between perl-style and python-style simulations.)
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/141>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#126: Simulation not deactivated during automatic cleanup
------------------------+---------------------------------------------------
Reporter: eschnett | Owner: mthomas
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
I had a simulation consisting of a single restart that finished, leaving
the output-0000-active link present. I then submitted a second restart.
While waiting in the queue, this link was still present; it was only
removed when the second restart actually run.
Submitting a second restart should only be possible if the simulation is
inactive, which means that this link is not present any more.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/126>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#116: Submitting an existing simulation on Kraken does not introduce the required
job dependency
------------------------+---------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
When I submit a simulation several times:
sim2 --remotemachine kraken submit parfiles/cptest.par 12 1:00:00
sim2 --remotemachine kraken submit parfiles/cptest.par
sim2 --remotemachine kraken submit parfiles/cptest.par
the queue state on Kraken is Q for all jobs. It should be Q for the first
and H for the second and third if the job dependency has been correctly
applied.
[hinder@kraken-pwd4 simulations]$ qstat -u hinder
nid00016: Kraken (UT/NICS Cray XT5)
Req'd Req'd Elap
Job ID Username Queue Jobname SessID NDS Tasks Memory Time S Time
866498.nid00016 hinder small cptest -- -- 12 -- 01:00 Q --
866499.nid00016 hinder small cptest -- -- 12 -- 24:00 Q --
866500.nid00016 hinder small cptest -- -- 12 -- 24:00 Q --
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/116>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#115: Remote submission does not pass on required options
------------------------+---------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
On my laptop, I execute a sim submit command with a --remotemachine
argument and a --config argument. The --config argument does not get
added to the remote command.
MacBook:llama ian$ sim2 --remotemachine damiana submit --config sim2
parfiles/cptest.par 4 1:00:00
DEBUG: Simfactory command: simfactory2/sim "--remotemachine" "damiana"
"submit" "--config" "sim2" "parfiles/cptest.par" "4" "1:00:00"
DEBUG: Version r992
The Simulation Factory: manage and submit cactus jobs
defs: /Users/ian/Cactus/llama/simfactory2/etc/defs.ini
defs.local: /Users/ian/Cactus/llama/simfactory2/etc/defs.local.ini
Managing for remote machine: damiana
Executing: /bin/bash -c "{ :; } && { :; } && ssh -Y ianhin@login-
damiana.aei.mpg.de \"/bin/bash -c '{ :; } &&
/home/ianhin/Cactus/llama/simfactory2/sim submit "parfiles/cptest.par" "4"
"1:00:00" '\""
Warning: No xauth data; using fake authentication data for X11 forwarding.
DEBUG: Simfactory command: /home/ianhin/Cactus/llama/simfactory2/sim
"submit" "parfiles/cptest.par" "4" "1:00:00"
DEBUG: Version r950
The Simulation Factory: manage and submit cactus jobs
defs: /home/ianhin/Cactus/llama/simfactory2/etc/defs.ini
defs.local: /home/ianhin/Cactus/llama/simfactory2/etc/defs.local.ini
Cactus Directory: /home/ianhin/Cactus/llama
Current Working directory does not match Cactus sourcetree, changing to
/home/ianhin/Cactus/llama
SimEnvironment.COMMAND: submit
Executing command: submit
Parfile: parfiles/cptest.par
Simulation Name: cptest
Procs: 4
Walltime: 1:00:00
Warning: simulation "cptest" does not exist or is not readable
Parameter file: /home/ianhin/Cactus/llama/parfiles/cptest.par
Configuration name not specified -- using default configuration "sim"
Configuration name not specified -- using default configuration "sim"
Warning: empty submit script for configuration sim
Error: empty/missing run script for configuration sim
Error 256 occured while executing command "/bin/bash -c "{ :; } && { :; }
&& ssh -Y ianhin(a)login-damiana.aei.mpg.de \"/bin/bash -c '{ :; } &&
/home/ianhin/Cactus/llama/simfactory2/sim submit "parfiles/cptest.par" "4"
"1:00:00" '\"""
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/115>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#111: Submitting an existing simulation does not take the walltime from the
previous restart
------------------------+---------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
I have submitted a simulation several times:
sim2 --remotemachine kraken submit parfiles/cptest.par 12 1:00:00
sim2 --remotemachine kraken submit parfiles/cptest.par
sim2 --remotemachine kraken submit parfiles/cptest.par
expecting the second and third jobs to inherit the walltime of the first
one. However, they appear in the qstat output on Kraken with 24 hours
(the default walltime).
[hinder@kraken-pwd4 simulations]$ qstat -u hinder
nid00016: Kraken (UT/NICS Cray XT5)
Req'd Req'd Elap
Job ID Username Queue Jobname SessID NDS Tasks
Memory Time S Time
-------------------- -------- -------- ---------------- ------ ----- -----
------ ----- - -----
866498.nid00016 hinder small cptest -- -- 12
-- 01:00 Q --
866499.nid00016 hinder small cptest -- -- 12
-- 24:00 Q --
866500.nid00016 hinder small cptest -- -- 12
-- 24:00 Q --
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/111>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#89: --remote execute doesn't cd into source directory
------------------------+---------------------------------------------------
Reporter: eschnett | Owner: mthomas
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
The command sim --remote numrel02 execute pwd doesn't cd into the source
directory before executing pwd; instead, it executed "pwd" in the home
directory.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/89>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#121: Should not be able to have two jobs in the queue for the same simulation at
the same time
------------------------+---------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
When a simulation is submitted, and there is a job in the queuing system
(either in the Q or the R state), it should not be possible to submit the
simulation again to run a concurrent job. Currently, a new job is created
in the Q state.
My approach would be to perform a qstat (or equivalent) and determine if
there is already a job in the queuing system from this simulation (by
checking the known job ids of the simulation against the output from
qstat). If there is no job in the queueing system, I would submit a new
job. If there are jobs in the queuing system, I would submit a "chained"
job to run when the last existing one is finished. This means that you
don't need new syntax to indicate that you want to chain. If it is
considered that this would be too confusing for new users, the decision of
whether to chain or abort when there is an existing job could be made
optional with a boolean option in the configuration file.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/121>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#165: Intel compiler complains wrongly about incompatible argument lists in
Fortran
----------------------+-----------------------------------------------------
Reporter: eschnett | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: Cactus | Version:
Keywords: |
----------------------+-----------------------------------------------------
On the CCT numrel workstations, the Intel compiler (with version and
options as defined in SimFactory) sometimes complains that argument lists
for calls in Fortran don't match. This usually happens to me in GRHydro.
The reason seems to be that the Intel compiler silently remembers the
prototype of every Fortran subroutine it encounters, and then compares
against this prototype (if it exists) when the subroutine is called. If a
subroutine's argument list changes, and if the caller is then recompiled
before the subroutine itself, then the compiler detects a mismatch
(because it compares to the old argument list) and aborts with an error.
The Intel compiler stores the argument lists together with module
information in the scratch directory. To solve this problem, one either
has to clean the configuration (make *-clean), or one has to manually
delete all outdated *.mod files belonging to this thorn (and possibly also
delete all *.o files of this thorn). A make *-clean is expensive, and the
manual solution requires in-depth knowledge of the problem.
Cactus should detect this problem and offer a solution, although I don't
know which. Maybe the solution could involve compiling files in a
different order by detecting the dependency between the caller and the
callee, and treating this in the same way as module definitions.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/165>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#118: EinsteinToolkit wiki should require a login for editing pages
-------------------------------------+--------------------------------------
Reporter: hinder | Owner:
Type: enhancement | Status: new
Priority: major | Milestone:
Component: EinsteinToolkit website | Version:
Keywords: |
-------------------------------------+--------------------------------------
The simfactory advanced tutorial at http://docs.einsteintoolkit.org/et-
docs/Simulation_Factory_Advanced_Tutorial was recently defaced. Since
this is now a problem in general for public-write-access services on the
web, I propose that the Einstein Toolkit wiki should require an account
for editing.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/118>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#175: Cactus webpage configuration option links are broken
---------------------+------------------------------------------------------
Reporter: bmundim | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Other | Version:
Keywords: |
---------------------+------------------------------------------------------
The lists of configuration options for Cactus are dead links
at http://cactuscode.org/download/configfiles/
I suspect they link to old simfactory files.
Thanks,
Bruno.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/175>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit