#1033: support a combined "map" file in CarpetIOHDF5
--------------------------+-------------------------------------------------
Reporter: rhaas | Owner: eschnett
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Carpet | Version:
Keywords: CarpetIOHDF5 |
--------------------------+-------------------------------------------------
recovering on different number of processes than was used to write a
checkpoint is painfully slow. Part of the reason seems to be that each
process essentially has to read all files to find out where each piece of
data it requires is located. The attached patch (not to be included in the
main code due to bad file formats and coding) enables CarpetIOHDF5 to read
all the information stored in the union of index files to from a single
file. This means (together with the other patches proposed today) that
CarpetIOHDF5 only ever opens those HDF5 files that are required to restore
the simulation on a given process. It significantly (factor > 4 where I
don't quite know how fast since the unpatched version ran out of walltime)
speeds up recovery with many more processors than wrote the files.
It also adds an optimization for CCTK_VarIndex calls inside CarpetIOHDF5
(which happens for every dataset in the file).
This is intended only as a proof of what might speed up recovery. A proper
implementation would need a more sensible file format. Two option seem
possible:
1) extend the index file format by a "filename" or "filenum" attribute to
each dataset and use a concatenation of all index files as the map file
2) define a custom hdf5 data type corresponding to the information in a
single patch_t, which would have mostly integer field plus two variable
length / enumerated ASCII fields (for the patch name, variable name)
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1033>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1737: upate CarpetHDF5 reader in VisIt
------------------------------+---------------------------------------------
Reporter: rhaas | Owner:
Type: enhancement | Status: new
Priority: minor | Milestone:
Component: Other | Version: development version
Keywords: VisIt CarpetHDF5 |
------------------------------+---------------------------------------------
Since the CarpetHDF5 reader
(https://svn.cactuscode.org/VizTools/CarpetHDF5) was included in VisIt's
main source code repo, we have not attempted to update that copy to the
most current version. The coding style in that branch matches the official
code but the patches do not apply on top of the official code due to
ordering issues as well as some minor new changes to code formatting in
the current (2.8.0) VisIt codebase.
Since then, I have collected a number of bugfixes/improvements that are
collected in various branches at https://bitbucket.org/rhaas80/carpethdf5
. It would be good to eventually try and get the for_VisIt branch
(https://bitbucket.org/rhaas80/carpethdf5/branch/for_VisIt) included in
the main source code repo again.
It contains some bug-fixes wrt how file metadata is cached, as well as
number of fixes to reduce memory footprint, number of open files (required
to work on large datasets on machines that limit the total number of open
files), some improvements to error reporting and robustness when dealing
with partially corrupted filesets as well as changes that allow a user to
combine multiple HDF5 files into a single "virtual" VisIt database using
.visit files.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1737>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1045: URL field should be optional
---------------------------+------------------------------------------------
Reporter: hinder | Owner: eric9
Type: defect | Status: new
Priority: minor | Milestone:
Component: GetComponents | Version:
Keywords: |
---------------------------+------------------------------------------------
In a CRL file, it should be possible to have only an AUTH_URL field and
omit the URL field, since it might be that there is no unauthenticated way
to access the repository (e.g. for private repositories). At the moment,
when I omit the URL field for a Git repository, the error message is:
{{{
Use of uninitialized value $git_repo in substitution (s///) at
./GetComponents line 589.
Use of uninitialized value $git_repo in substitution (s///) at
./GetComponents line 590.
Use of uninitialized value $rec{"GIT_REPO"} in substitution (s///) at
./GetComponents line 593.
Use of uninitialized value $rec{"GIT_REPO"} in substitution (s///) at
./GetComponents line 594.
...
}}}
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1045>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1772: Simfactory: potentially serious problem with CACHE directory in the
simulations directory
------------------------+---------------------------------------------------
Reporter: bmundim | Owner:
Type: defect | Status: new
Priority: critical | Milestone: ET_2015_05
Component: SimFactory | Version: development version
Keywords: CACHE |
------------------------+---------------------------------------------------
Is the directory CACHE in the simulation directory really necessary? We
are talking about executables with at most 400MB of size, which is nothing
compared to current HPC storage systems.
I think I might have found a design flaw on simfactory use of CACHE
directory which can go unnoticed until it is too late with potential loss
of thousands of SUs. Suppose we have the following situation:
1) We build a configuration A and send a simulation A1 with with parameter
file 1. So simfactory copies the executable from configuration A to
simulation A1 simfactory directory and creates a symlink from
/scratch/simulations/CACHE/exe/cactus_A to
/scratch/simulations/A1/SIMFACTORY/exe/cactus_A.
2) We then create a new simulation A2 with a different parameter file 2.
This time simfactory symlink the simulation executable
/scratch/simulations/A2/SIMFACTORY/exe/cactus_A to the cached one
/scratch/simulations/CACHE/exe/cactus_A.
3) After a few days (or restarts) of simulations A1 and A2, you come up
with a better idea/fix/new parameter which requires to recompile your
configuration A. Note that we don't want to build a new configuration from
scratch since cactus configurations consume both a lot of time and space
to build. So you rebuild your configuration A and its executable cactus_A
is updated.
4) Let's say now we submit the updated configuration with the same
parameter file 2 in order to test your new idea/fix/parameter and compare
it with the simulation A2, which is still running and have a few extra
restarts to completion. Call this simulation A2_updated. Simfactory then
copy the new updated executable cactus_A from the Cactus/exe/cactus_A to
the simulation directory
/scratch/simulations/A2_updated/SIMFACTORY/exe/cactus_A *and* update the
CACHE symlink to that new simulation directory, ie:
$ cd /scratch/simulations/CACHE/exe
$ ls -l cactus_A
cactus_A ->
../../../../scratch/simulations/A2_updated/SIMFACTORY/exe/cactus_A
5) The problem: now my simulation A2 restarts are compromised with a new
executable. Remember that that simulation executable is actually a symlink
to the one in the CACHE directory, which has just been updated.
I think this whole cache directory intermediate step introduces
unnecessary complexity for the user to track; it is really unnecessary and
in my opinion not a good design choice. I would vote to eliminate it from
simfactory completely as soon as possible, ideally even for this release.
Just use one copy of the executable from cactus/exe to
simulation/SIMFACTORY/exe and that's it. This is all we need to have that
simulation and future ones running consistently with the same executable.
Thanks!
PS: I have actually noticed this issue on Hershel release (there is no
option pointing to Hershel release on trac). I am working on tests for
development version to confirm this issue, but give simfactory commits I
believe it is still there.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1772>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1749: avoid creating temporary links to files in Formaline
-----------------------------------+----------------------------------------
Reporter: rhaas | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version: development version
Keywords: Formaline |
-----------------------------------+----------------------------------------
the patch in https://bitbucket.org/cactuscode/cactusutils/branch/rhaas
%2Fupdate-index changes the way Formaline populates the configjar git
repository with changed files. Instead of making a hard linked copy of
every file in the thorn, it used plumbing commands git-update-index to add
them to the index. This deals gracefully with files in symbolically linked
directories (but will dereference a symbolic link if the file *itself* is
a symgolic link) and also works when CACTUS_CONFIGS_DIR is not on the same
file system as the source tree (eg Cactus is in $HOME which has a small
quota but is backed up and $CACTUS_CONFIGS_DIR is in $SCRATCH which is not
backed up).
Not being able to have $CACTUS_CONFIGS_DIR on a different file system than
Cactus is a minor bug, in particular since this method is described in the
UserGuide.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1749>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#334: unsuccessful qsub not recognized / submit succeeds for finished simulation
------------------------+---------------------------------------------------
Reporter: knarf | Owner: mthomas
Type: defect | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
python version:
I submitted a simulation 'sim submit' but the corresponding qsub failed
due to wrong numbers of procs/node (philip cluster). I changed the number
given on the command line and did a 'sim submit' again, this time
successful. Several things happend which I think could be done better:
- the unsuccessful qsub was not detected during the new submit - it
attempted a restart and didn't simply clean
the unsuccessful submit
- when trying the restart, it went ahead and queued the job, but this
later failed when run with "cannot rerun a restart that has been
finished". This could have been caught earlier - without the wait time in
the queue.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/334>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1790: ADMMass: Properly distinguish between int and CCTK_INT
-----------------------------------+----------------------------------------
Reporter: eschnett | Owner:
Type: defect | Status: new
Priority: minor | Milestone:
Component: EinsteinToolkit thorn | Version: development version
Keywords: |
-----------------------------------+----------------------------------------
Patch attached.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1790>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1627: Merge rewrite branch of McLachlan
-----------------------------------+----------------------------------------
Reporter: eschnett | Owner:
Type: enhancement | Status: new
Priority: major | Milestone:
Component: EinsteinToolkit thorn | Version: development version
Keywords: |
-----------------------------------+----------------------------------------
We should merge the "rewrite" branch of McLachlan.
To do before this merge:
- Remove helper thorn
- Add backward compatibility layer for parameters
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1627>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#649: simfactory should create the "simulations" directory
------------------------+---------------------------------------------------
Reporter: rhaas | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version:
Keywords: |
------------------------+---------------------------------------------------
when following the new users tutorial one gets:
{{{
[rhaas@qb1 Cactus]$ ./simfactory/bin/sim submit static_tov
--parfile=par/static_tov.par --procs=32 --walltime=8:0:0
Parameter file: /home/rhaas/Cactus/par/static_tov.par
Error: could not access simulation base directory
/scratch/rhaas/simulations for reading and writing
Aborting Simfactory.
[137450 refs]
}}}
This happens each time I set up a fresh simfactory on a machine. It might
be useful if simfactory would create the simulation base directory if it
does not exist.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/649>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#1645: simfactory should create exe directory
------------------------+---------------------------------------------------
Reporter: rhaas | Owner: eschnett
Type: defect | Status: new
Priority: minor | Milestone:
Component: SimFactory | Version: development version
Keywords: |
------------------------+---------------------------------------------------
Using virtual (--virtual) configuration one can wrap a Cactus config (or
at least the parts that simfactory cares about) around an existing
executable. Simfactory then copies the executable to exe/cactus_sim (or
whatever), however it assumes that exe already exists. Since exe is not
part of the Cactus checkout, simfactory needs to itself ensure that it
exists.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/1645>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit