#63: create separate tracs on trac.cactuscode.org and trac.cfdtoolkit.org
--------------------+-------------------------------------------------------
Reporter: knarf | Owner:
Type: task | Status: new
Priority: minor | Milestone:
Component: Cactus | Version:
Keywords: |
--------------------+-------------------------------------------------------
We need separate tracs on trac.cactuscode.org and trac.cfdtoolkit.org,
including ssl certificates and the complete rest of the setup.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/63>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#233: Unable to complete an ET checkout without SVN network errors
--------------------+-------------------------------------------------------
Reporter: hinder | Owner:
Type: defect | Status: new
Priority: major | Milestone:
Component: Other | Version:
Keywords: |
--------------------+-------------------------------------------------------
I am running an automated build and test system every night which attempts
to check out the Einstein Toolkit thornlist. Every night, I get error
messages from at least one thorn of the form:
Checking out module: CactusNumerical/Cartoon2D
from repository:
http://svn.cactuscode.org/arrangements/CactusNumerical/Cartoon2D/trunk
into: ./arrangements
svn: REPORT of '/arrangements/CactusNumerical/Cartoon2D/!svn/vcc/default':
Could not read response body: connection was closed by server.
(http://svn.cactuscode.org)
Typically, three or more thorns will exhibit this problem. All the
following repositories have exhibited this problem since 11th January
2011:
CactusArchive/ADM
CactusElliptic/EllBase
CactusNumerical/Cartoon2D
CactusNumerical/Dissipation
CactusNumerical/InterpToArray
CactusNumerical/RotatingSymmetry180
CactusNumerical/RotatingSymmetry90
EinsteinInitialData/Exact
EinsteinInitialData/IDAnalyticBH
EinsteinInitialData/TOVSolver
EinsteinInitialData/TwoPunctures
LSUThorns/QuasiLocalMeasures
manifest
simfactory
I assume that there is some problem with the SVN server or our network
connection to it that is making it so unreliable. Am I the only one?
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/233>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#317: Unproper bevaviour of sim submit when changing the parfile
----------------------------------------------+-----------------------------
Reporter: alexander.beck-ratzka@… | Owner: mthomas
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version: ET_2010_11
Keywords: |
----------------------------------------------+-----------------------------
If I make changes to the parfile of my simulation, and call then
simfactory submit with the option --parfile=newpar, simfactory does not
take this parfile. Instead of this, simfactory uses an old parfile from
the old output-xxxx directory in my simualations direrctory.
This usage of the old parfile might be wished. However, if this is the
case, simfactory must complain if invoked with the option "parfile=...".
But simfactory does not complain.
In my case, the parfile has been modfied, but I am using the same name.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/317>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#123: Run a syntax checker in a pre-commit hook
------------------------+---------------------------------------------------
Reporter: eschnett | Owner: mthomas
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
I am getting annoyed by the frequent syntax errors in SimFactory that used
to be caught by Perl and are now not caught by Python. It would be ideal
if syntax checking could be built into SimFactory, so that it executes
automatically during startup if that isn't too slow.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/123>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#124: Implement a testing mechanism for simfactory
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: enhancement | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
-------------------------+--------------------------------------------------
Currently there is no easy way to check that a change to simfactory hasn't
broken critical functionality. A system could be introduced to run
through each command and check that it does the right thing. It won't be
possible to check every single combination of possible options, but the
most commonly used options should be tested.
A first attempt could involve testing that the commands work locally on
the current machine. This would involve submitting jobs and testing that
they have the expected behaviour. On busy machines, this means that the
tests could potentially take a long time to complete.
Remote submission could be tested by running these tests from a central
location (e.g. the developer's workstation or laptop).
We could also include a python correctness checking step in these tests
(e.g. using pychecker or pylint).
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/124>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#180: Reduce disk space used by checkpoints
-------------------------+--------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: enhancement | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
-------------------------+--------------------------------------------------
Currently simfactory stores the checkpoints for each restart in their own
directory. This means that the Cactus mechanism for deleting all but the
most recent N checkpoints does not see the previous checkpoints. This
means that you can easily run out of quota space when doing very long
simulations.
One solution to this would be for SimFactory to store the checkpoints in a
directory above the output-NNNN directories and make the current
checkpoints directories under output-NNNN symbolic links to the common
directory. This way, all restarts would see the same directory for
checkpoint files, and Cactus could clean up the old checkpoints. The
current hardlinking mechanism would not be required any more. This
solution might be undesirable because it means that each restart is no
longer independent.
Another solution would be to give simfactory an option to delete old
checkpoint files from previous restarts when a job starts. This would
duplicate the functionality already available in Cactus.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/180>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#285: python simfactory submit does not remove active link after job execution
has finished
----------------------------------------------+-----------------------------
Reporter: alexander.beck-ratzka@… | Type: defect
Status: new | Priority: major
Milestone: | Component: Other
Version: | Keywords:
----------------------------------------------+-----------------------------
The python version of simfactory does not remove the active link
output-xxxx-active
in the simulation directory, after a job has finished.
As a consequence the next submit fails with the error message, that more
then one active link has been found. Removing the link manually enables a
poper submit of the next job. The problem occurs on a lustre file system.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/285>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#113: Simfactory should be able to run the Cactus testsuites
-------------------------+--------------------------------------------------
Reporter: knarf | Owner: mthomas
Type: enhancement | Status: new
Priority: critical | Milestone: ET_2011_06
Component: SimFactory | Version:
Keywords: |
-------------------------+--------------------------------------------------
... in the python version, before the next ET release. The scripts at
http://einsteintoolkit.org/release-info/ might help with implementing
that.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/113>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#316: Checkpoint recovery nonfunctional
------------------------+---------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: defect | Status: new
Priority: blocker | Milestone:
Component: SimFactory | Version:
Keywords: regression |
------------------------+---------------------------------------------------
Checkpoint recovery is nonfunctional in SimFactory 2 (it has broken since
it was last fixed in ticket #60).
Using the attached parameter file, I submit a simulation on Datura:
simfactory2/bin/sim --machine datura --config sim2_datura create-submit
parfiles/cptest.par 12 1:00:00
This parameter file terminates the Cactus run after 1 minute and dumps a
checkpoint file. I then manually remove the output-0000-active symlink,
as the automatic cleanup in the main() function is cleaning up restarts
that are attempting to run, so I have disabled it, and manual cleanup
doesn't work (see ticket #315).
I then resubmit the simulation
simfactory2/bin/sim --machine datura submit parfiles/cptest.par
and observe that the checkpoint files from the first restart are never
hardlinked into the output directory. The job does not recover, and
instead starts from initial data.
Log file is attached.
Looking at the code, it appears that the checkpoint linking is conditional
on the from-restart-id parameter being passed to simfactory, which I think
is something to do with job-chaining. I can't see anywhere in the code
which sets this option, so this is probably why the linking is not
happening.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/316>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#304: SimFactory should not clean-up so aggressively
------------------------+---------------------------------------------------
Reporter: hinder | Owner: mthomas
Type: defect | Status: new
Priority: major | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
As far as I can tell, SimFactory 2 currently cleans up all simulations
every time it is run. This is not scalable - if you are running with a
slow production filesystem, just statting all the simulations could take a
very long time. Similarly, if you do a sim sync, it currently cleans up
all the simulations on your local machine, as well as the remote machine.
Running sim --help even cleans up all your simulations!
This behaviour is very counter-intuitive (certainly not what a user would
expect) and not suitable for production systems. I propose that
simfactory should only clean up the simulation which is being referred to
in the current command.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/304>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit