#180: Reduce disk space used by checkpoints
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: enhancement | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
-------------------------+--------------------------------------------------
Currently simfactory stores the checkpoints for each restart in their own
directory. This means that the Cactus mechanism for deleting all but the
most recent N checkpoints does not see the previous checkpoints. This
means that you can easily run out of quota space when doing very long
simulations.
One solution to this would be for SimFactory to store the checkpoints in a
directory above the output-NNNN directories and make the current
checkpoints directories under output-NNNN symbolic links to the common
directory. This way, all restarts would see the same directory for
checkpoint files, and Cactus could clean up the old checkpoints. The
current hardlinking mechanism would not be required any more. This
solution might be undesirable because it means that each restart is no
longer independent.
Another solution would be to give simfactory an option to delete old
checkpoint files from previous restarts when a job starts. This would
duplicate the functionality already available in Cactus.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/180>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#285: python simfactory submit does not remove active link after job execution
has finished
----------------------------------------------+-----------------------------
Reporter: alexander.beck-ratzka@… | Type: defect
Status: new | Priority: major
Milestone: | Component: Other
Version: | Keywords:
----------------------------------------------+-----------------------------
The python version of simfactory does not remove the active link
output-xxxx-active
in the simulation directory, after a job has finished.
As a consequence the next submit fails with the error message, that more
then one active link has been found. Removing the link manually enables a
poper submit of the next job. The problem occurs on a lustre file system.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/285>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#461: Filter rules error when using rsync 2.6.9
------------------------+---------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: defect | Status: new
Priority: blocker | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
When I use rsync 2.6.9 (the default for Mac OS 10.6) on my local machine,
I get the error
invalid modifier sequence at 'p' in filter rule: -p _darcs
when running sim sync. Probably this is an rsync 3 feature.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/461>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#469: Improve default directory locations
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: eschnett
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Keywords: |
-------------------------+--------------------------------------------------
It is usually desirable to set the user's source base directory to be in
their home directory. On many systems, this can be inferred from the user
name; e.g. /home/@USER@. However, on systems with very large numbers of
users, the home directory is sometimes not predictable in this way. For
example, /home1/00915/@USER@ (this is on LoneStar), where the numbers are
different for different users. This also happens on Kraken (see ticket
#382). Having to specify these details manually for each machine that you
want to use is an extra step in configuring SimFactory which we would like
to avoid.
Option 1: Provide a @HOME@ definition which is determined automatically
from the user's home directory.
This would be possible, but it only solves the problem for the home
directory. On some systems (again, LoneStar and Kraken), the home
directory might not be large enough to use as a source base directory, and
instead you might want to use the "work" directory instead (this is the
current simfactory setting for LoneStar). This directory also suffers
from the same unpredictable numbering problems. As the home directory is
available in the HOME environment variable, the work directory is
available in the WORK environment variable.
Option 2: Provide a general method for accessing environment variables and
storing them in definitions. For example, you could write @ENV(HOME)@ or
@ENV(WORK)@. This requires the environment to have been correctly set up
(for remote operation, this requires a login shell, but I think we have
that already). In fact it looks like there is preliminary code in
simsubs.py, but it is incomplete and is not called from anywhere.
I would favour option 2. Thoughts?
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/469>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#381: Cannot log in to Kraken using default simfactory configuration
------------------------+---------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: defect | Status: new
Priority: major | Milestone: ET_2011_05
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
SimFactory uses the gsissh method to connect to Kraken. This requires a
"proxy" to be created locally via the myproxy-logon command. There is
currently logic in Perl simfactory's mdb.pm file to automatically run
myproxy-logon and prompt for the passphrase.
This does not work for me. I get the error
Failed to receive credentials.
ERROR from myproxy-server (myproxy.teragrid.org):
PAM authentication failed: Permission denied
(see below for the full output). If I remove the localsshsetup key, I can
run the myproxy-logon command manually and then connect using simfactory.
Looking at the full output below, it seems that it is trying to run
myproxy-logon with my local username (ian) instead of the remote one
(hinder). If I replace the @USER@ in the myproxy-logon command with my
Kraken username, everything works as expected.
MacBook:etrelease ian$ sim login kraken
Simulation Factory:
Executing: {
mkdir -p /Users/ian/Cactus/etrelease/.globus &&
: >> /Users/ian/Cactus/etrelease/.globus/proxy-teragrid &&
chmod go-rwx /Users/ian/Cactus/etrelease/.globus/proxy-teragrid &&
export X509_USER_PROXY=/Users/ian/Cactus/etrelease/.globus/proxy-teragrid
&&
: mkdir -p /Users/ian/Cactus/etrelease/.globus/certificates-teragrid &&
: export X509_CERT_DIR=/Users/ian/Cactus/etrelease/.globus/certificates-
teragrid &&
{
{
grid-proxy-info -issuer -file /Users/ian/Cactus/etrelease/.globus
/proxy-teragrid 2>/dev/null | grep "^/C=US/O=National Center for
Supercomputing Applications/" >/dev/null 2>/dev/null &&
test $(grid-proxy-info -timeleft -file
/Users/ian/Cactus/etrelease/.globus/proxy-teragrid 2>/dev/null) -gt 0
2>/dev/null
} || {
{ grid-proxy-destroy 2>/dev/null || true; } &&
myproxy-logon -p 7514 -s myproxy.teragrid.org -T -l ian -o
/Users/ian/Cactus/etrelease/.globus/proxy-teragrid
}
} &&
{ globus-update-certificate-dir > /dev/null 2>&1 || true; }; } && gsissh
-t hinder(a)kraken-gsi2.nics.teragrid.org '{ source /etc/profile; } && cd
/nics/b/home/hinder/Cactus/etrelease && $SHELL -l'
Enter MyProxy pass phrase:
Failed to receive credentials.
ERROR from myproxy-server (myproxy.teragrid.org):
PAM authentication failed: Permission denied
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/381>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#113: Simfactory should be able to run the Cactus testsuites
-------------------------+--------------------------------------------------
Reporter: knarf | Owner: mthomas
Type: enhancement | Status: new
Priority: critical | Milestone: ET_2011_06
Component: SimFactory | Version:
Keywords: |
-------------------------+--------------------------------------------------
... in the python version, before the next ET release. The scripts at
http://einsteintoolkit.org/release-info/ might help with implementing
that.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/113>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#466: --reconfig seems to change thorn list
------------------------+---------------------------------------------------
Reporter: eschnett | Owner: eschnett
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
The option --reconfig (without also giving a thorn list) seems to change
the thorn list.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/466>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#464: create-submit without simulation name gives bad error message
------------------------+---------------------------------------------------
Reporter: eschnett | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
Calling create-submit without specifying a simulation name does not give a
legible error message:
$ ./bin/sim create-submit --debug
Info: Simfactory command: ./bin/../simfactory/lib/sim.py "create-submit" "
--debug"
Info: Version 1371M
The Simulation Factory: Manage Cactus simulations
Info: defs: /Users/eschnett/EinsteinToolkit-hg-
vanilla/simfactory/etc/defs.ini
Info: defs.local: /Users/eschnett/EinsteinToolkit-hg-
vanilla/simfactory/etc/defs.local.ini
Info: Cactus Directory: /Users/eschnett/EinsteinToolkit-hg-vanilla
Info: simenv.COMMAND: create-submit
Info: Executing command: create_submit
Traceback (most recent call last):
File "./bin/../simfactory/lib/sim.py", line 147, in <module>
main()
File "./bin/../simfactory/lib/sim.py", line 143, in main
CommandDispatch()
File "./bin/../simfactory/lib/sim.py", line 105, in CommandDispatch
module.main()
File "/Users/eschnett/EinsteinToolkit-hg-vanilla/simfactory/lib/sim-
manage.py", line 345, in main
CommandDispatch()
File "/Users/eschnett/EinsteinToolkit-hg-vanilla/simfactory/lib/sim-
manage.py", line 324, in CommandDispatch
exec("command_%s()" % command)
File "<string>", line 1, in <module>
File "/Users/eschnett/EinsteinToolkit-hg-vanilla/simfactory/lib/sim-
manage.py", line 125, in command_create_submit
simulationName = simenv.OptionsManager.args.pop(0)
IndexError: pop from empty list
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/464>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#455: Sort link directories so that system directories come last
----------------------+-----------------------------------------------------
Reporter: eschnett | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Cactus | Version:
Keywords: |
----------------------+-----------------------------------------------------
The enclosed patch ensures (makes it more likely) that external libraries
provided by Cactus ("ExternalLibraries") supersede system libraries. The
underlying problem is that all library search paths specified via -L are
concatenated on the final link line, and if different versions of the same
library exist, it is then difficult to ensure that that right one is used.
The enclosed patch sorts the directories in final link line such that
system directories (in particular /opt/local) come last, so that Cactus's
directories come first.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/455>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit