Present: Frank, Roland, Ian, Erik, Steve, Zach, Peter
GRHydro update: * GRHydro is seeing a major rewrite with most F90 routines being re-implemented as C++ * in this process seldom (or completely) unused features are being removed * suggestion is to announce changes (rewrite and lost functionality) on mailing list then drop old code * Roland to make a list of changed/dropped functionality
Git transition: * we will combine related thorns into common repositories * keep some thorns (ExternalLibraries that contain tarballs) in svn * Erik to create a proposed list of repositories and thorns in them on the wiki
Timelevels: * changes to MoL to avoid having to count evolved variables * desire to have similar auto-counting in Carpet to determine number of timelevels required for interpolation in time * currently seems as if Carpet does not have enough information early enough to allocate those timelevels * suggestion (Erik): "register" grid functions with Carpet that same way they are registered with MoL if interpolation in time will be required. This will be a subset of the grid functions that do *not* state prolongation_type = None
Yours, Roland
On Tue, Mar 18, 2014 at 09:17:38AM -0700, Roland Haas wrote:
Git transition:
- we will combine related thorns into common repositories
- keep some thorns (ExternalLibraries that contain tarballs) in svn
- Erik to create a proposed list of repositories and thorns in them on
the wiki
Actually I volunteered, and you can view the list here:
https://docs.einsteintoolkit.org/et-docs/Repository_transition
See this as proposal. This essentially puts every thorn into its own repository, but combines:
- PITTNullCode - GRHydro - CactusExamples - CactusTest - CactusWave
It also leaves out ExternalLibraries, at least for now. When we are going to move to the proposed mechanism, these repositories would need to be created from scratch anyway.
Frank
On Mon, Mar 24, 2014 at 9:50 AM, Frank Loeffler knarf@cct.lsu.edu wrote:
See this as proposal. This essentially puts every thorn into its own repository, but combines:
- PITTNullCode
- GRHydro
- CactusExamples
- CactusTest
- CactusWave
It also leaves out ExternalLibraries, at least for now. When we are going to move to the proposed mechanism, these repositories would need to be created from scratch anyway.
It's not important right now since the ExternalLibraries won't be transitioning yet, but while I think of it, it would be possible to maintain the history rather than creating the repository from scratch. This could be done using git filter-branch to remove the tarballs from the history and put them in a submodule instead. Checking out the out the thorn would then not get all the tarballs by default. Of course, if a tarball is needed then the entire history would have to be downloaded. However, if the ExternalLibraries eventually move to a system where the tarball is not included but instead fetched on demand, then the large history-of-tarballs repository would only have to be fetched when checking out an old version of the external library.
Hi
On Mon, Mar 24, 2014 at 11:29:32AM -0400, Barry Wardell wrote:
It's not important right now since the ExternalLibraries won't be transitioning yet, but while I think of it, it would be possible to maintain the history rather than creating the repository from scratch.
I am sure it can be done. The question here is whether it is worth it. I don't think the history of these repositories is really worth transitioning - the svn repositories will stay around anyway.
included but instead fetched on demand, then the large history-of-tarballs repository would only have to be fetched when checking out an old version of the external library.
Which would mean git would need to be installed wherever you compile Cactus. We cannot assume that, and I don't really want to build git first.
Frank
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
Hello all,
Which would mean git would need to be installed wherever you compile Cactus. We cannot assume that, and I don't really want to build git first.
The ExternalLibraries are ET not Cactus, right? We already agreed that the ET will transition to git (and in the process apparently also transition Cactus* to git). So we already require git to be present I think.
I am also not aware of a machine that is in eg simfactory's mdb that does not contain (some) version of git.
Finally I think rather than checking out a repository of tarballs I'd try and use eg bitbuckets' option to access a git repository via http(s) to download only the single tarball of interest (or even have a svn repository just for the tarballs or download the tarballs from the developer site [assuming that most projects for which we'd need tarballs are longer lived than the ET]). Possibly a --shallow git repository for the tarball can also be used (though that still keeps history on the local disk when one updates).
Yours, Roland
- -- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
On Mon, Mar 24, 2014 at 01:05:25PM -0700, Roland Haas wrote:
The ExternalLibraries are ET not Cactus, right? We already agreed that the ET will transition to git (and in the process apparently also transition Cactus* to git). So we already require git to be present I think.
They are hosted by Cactus, but are obviously not requiring LGPL. We do not require git to be present where you compile. You require it only where you check things out. You then use rsync/simfactory to copy your tree wherever you want it.
I am also not aware of a machine that is in eg simfactory's mdb that does not contain (some) version of git.
You would be surprised. :) But then, maybe not.
We agreed to download the tarball, only if needed, at runtime. This should work reasonably well using either wget or curl on essentially all systems. Whether these are then stored verbatim, in svn or git is a separate issue, and does not matter here I think. We don't have the machinery in place to do this anyway right now, so I suggest to not mix this with the general svn->git transition, and leave the ExternalLibraries where they are, for now - or make sure we get the scripts in place that would handle the downloading ect. first.
Frank
On 24 Mar 2014, at 21:10, Frank Loeffler knarf@cct.lsu.edu wrote:
On Mon, Mar 24, 2014 at 01:05:25PM -0700, Roland Haas wrote:
The ExternalLibraries are ET not Cactus, right? We already agreed that the ET will transition to git (and in the process apparently also transition Cactus* to git). So we already require git to be present I think.
They are hosted by Cactus, but are obviously not requiring LGPL. We do not require git to be present where you compile. You require it only where you check things out. You then use rsync/simfactory to copy your tree wherever you want it.
I am also not aware of a machine that is in eg simfactory's mdb that does not contain (some) version of git.
You would be surprised. :) But then, maybe not.
We agreed to download the tarball, only if needed, at runtime. This should work reasonably well using either wget or curl on essentially all systems.
I don't remember an agreement on this; while it is my preferred solution (along with having the possibility of pre-caching the tarballs with a single command), I don't think we had agreement from everyone on this.
On Mon, Mar 24, 2014 at 11:38:32PM +0100, Ian Hinder wrote:
I don't remember an agreement on this; while it is my preferred solution (along with having the possibility of pre-caching the tarballs with a single command), I don't think we had agreement from everyone on this.
Ok, then we need to discuss this, either here or/and at the call next week.
Frank
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
Hello Frank, Ian,
Ok, then we need to discuss this, either here or/and at the call next week.
According to the minutes from 2013-12-16 a decision was reached: - --8<-- External libraries size issues: * all want to reduce link time * executable size, CactusJar size reduction welcomed by all * Ian would like to reduce overall repository size for machines with tight quotas on $HOME * will implement the following scheme using Boost as a testing ground: ** ExternalLibraries no longer contain source tarball ** if at build time they detect that they need to compile from source, they will download the tarball (if not already there) into a Cactus-tree wide cache directory ** these source tarballs will not be part of Formaline tarball, however the actual ExtenalLibraries thorn (ie patches, configures.sh, an thus the exact URL used to download the tarball) source is part of the Formaline tarball. This ensures that the code can be recompiled in the future using only information from the Formaline tarball - --8<-- This is on the mailing list archive http://lists.einsteintoolkit.org/pipermail/users/2013-December/003383.html (is there any progress on ticket 719 https://trac.einsteintoolkit.org/ticket/719 ?).
What is not in the notes is that Ian also desired a way to force it to download so that for cluster where no network connection (or git) is available at build time (or preparing for a flight) one can force download on ones laptop then use simfactory to sync the tarballs. It was actually proposed to try this for a thorn to see how this works after which we got sidetracked discussing boost's build system.
In the end Ian and Frank actually agreed on a scheme (mostly following Ian's suggestions). Someone would have to test it though. Otherwise I would not want to open that particular can of worms again.
Yours, Roland
- -- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
On Mon, Mar 24, 2014 at 04:07:35PM -0700, Roland Haas wrote:
According to the minutes from 2013-12-16 a decision was reached:
Thanks for digging this out. This is what I remembered.
What is not in the notes is that Ian also desired a way to force it to download so that for cluster where no network connection (or git) is available at build time (or preparing for a flight) one can force download on ones laptop then use simfactory to sync the tarballs.
I think we mentioned at least in the call that such a possibility should exist - an option to download all possibly required tarballs.
Frank
On 25 Mar 2014, at 14:53, Frank Loeffler knarf@cct.lsu.edu wrote:
On Mon, Mar 24, 2014 at 04:07:35PM -0700, Roland Haas wrote:
According to the minutes from 2013-12-16 a decision was reached:
Thanks for digging this out. This is what I remembered.
What is not in the notes is that Ian also desired a way to force it to download so that for cluster where no network connection (or git) is available at build time (or preparing for a flight) one can force download on ones laptop then use simfactory to sync the tarballs.
I think we mentioned at least in the call that such a possibility should exist - an option to download all possibly required tarballs.
The obvious place to store the URL for the tarball is in the configuration.ccl file. What is the recommended procedure for reading entries from this file outside of the build system? It would have been nice to find that the Perl files in lib/sbin were Perl modules which could be used outside the build system, but I was disappointed to find this was not the case. The file is parsed by lib/sbin/ConfigurationParser.pl.
On Mar 25, 2014, at 10:30 , Ian Hinder ian.hinder@aei.mpg.de wrote:
On 25 Mar 2014, at 14:53, Frank Loeffler knarf@cct.lsu.edu wrote:
On Mon, Mar 24, 2014 at 04:07:35PM -0700, Roland Haas wrote:
According to the minutes from 2013-12-16 a decision was reached:
Thanks for digging this out. This is what I remembered.
What is not in the notes is that Ian also desired a way to force it to download so that for cluster where no network connection (or git) is available at build time (or preparing for a flight) one can force download on ones laptop then use simfactory to sync the tarballs.
I think we mentioned at least in the call that such a possibility should exist - an option to download all possibly required tarballs.
The obvious place to store the URL for the tarball is in the configuration.ccl file. What is the recommended procedure for reading entries from this file outside of the build system? It would have been nice to find that the Perl files in lib/sbin were Perl modules which could be used outside the build system, but I was disappointed to find this was not the case. The file is parsed by lib/sbin/ConfigurationParser.pl.
I think we want to move away from building external libraries as part of the Cactus build; instead, we want to build them ahead of time, e.g. via Simfactory. This has several advantages, such as e.g. that a "make clean" doesn't require rebuilding the libraries, and that several Cactus configurations can use the same external libraries, and that one can even have one power user build the libraries on a system while all others simply use them.
One additional advantage is that building external libraries from within Cactus is actually quite difficult, since the Cactus compiler options are often "strange" and need to be "cleaned" before one can use them. The respective shell code is arcane, and often breaks between different system. Building external libraries on their own, independent of Cactus, is much simpler.
I know this because I implemented this in Simfactory3, and I definitively think this is the way to go. The current build recipes use Clang (not gcc, and not the system's "standard" compiler) for building. I did this because I was interested in C++11, and most system compilers (Intel, PGI, older versions of GCC, current Nvidia) don't support this. However, in the interest of backward compatibility we should probably also provide build recipes for (say) an older version of GCC that is supported by all compilers.
-erik
On Tue, Mar 25, 2014 at 10:58:31AM -0400, Erik Schnetter wrote:
I think we want to move away from building external libraries as part of the Cactus build; instead, we want to build them ahead of time, e.g. via Simfactory. This has several advantages, such as e.g. that a "make clean" doesn't require rebuilding the libraries, and that several Cactus configurations can use the same external libraries, and that one can even have one power user build the libraries on a system while all others simply use them.
I think this is a good idea.
I know this because I implemented this in Simfactory3, and I definitively think this is the way to go. The current build recipes use Clang (not gcc, and not the system's "standard" compiler) for building. I did this because I was interested in C++11, and most system compilers (Intel, PGI, older versions of GCC, current Nvidia) don't support this. However, in the interest of backward compatibility we should probably also provide build recipes for (say) an older version of GCC that is supported by all compilers.
We probably would need that, and I myself would want that too. gcc is a pseudo-standard we have to support.
Frank
On Mar 25, 2014, at 11:03 , Frank Löffler knarf@cct.lsu.edu wrote:
On Tue, Mar 25, 2014 at 10:58:31AM -0400, Erik Schnetter wrote:
I think we want to move away from building external libraries as part of the Cactus build; instead, we want to build them ahead of time, e.g. via Simfactory. This has several advantages, such as e.g. that a "make clean" doesn't require rebuilding the libraries, and that several Cactus configurations can use the same external libraries, and that one can even have one power user build the libraries on a system while all others simply use them.
I think this is a good idea.
I know this because I implemented this in Simfactory3, and I definitively think this is the way to go. The current build recipes use Clang (not gcc, and not the system's "standard" compiler) for building. I did this because I was interested in C++11, and most system compilers (Intel, PGI, older versions of GCC, current Nvidia) don't support this. However, in the interest of backward compatibility we should probably also provide build recipes for (say) an older version of GCC that is supported by all compilers.
We probably would need that, and I myself would want that too. gcc is a pseudo-standard we have to support.
I didn't mean that we support building BY gcc; what I meant was that we build the external libraries WITH gcc, so that they work with any other compiler.
When building libraries, it is important to use compatible libstdc++ libraries, and this is a choice made by the C++ compiler. Thus, e.g. "building OpenMPI with GCC" and "building OpenMPI with Intel" should not make a difference at all; what is necessary is to use the same version of libstdc++ when building OpenMPI. Typically, there is one version of libstdc++ installed on the system, and this version may be old, and if you install GCC yourself you get a newer version, and the Intel compiler picks one of these versions...
Thus my suggestion is to install GCC ourselves, using an outdated version such as 4.6 that is actually supported by Intel and Nvidia, and then using its libstdc++ to build all other external libraries, as well as Cactus. Whether we then use GCC or Intel or PGI or Clang to actually compile C++ code then doesn't matter much; the generated code will be compatible.
That's what I've been doing in Simfactory3, and it seems to be working nicely. Except that some package need a lot of arm twisting to actually use a particular version of libstdc++, and that some system administrators use setups that try hard to override anything you're trying to do, but if you combine cannons (for the sparrows) and sledgehammers and a certain lack of ruth, you can do it.
I should give an overview over this system next Monday.
-erik
Hello all,
I know this because I implemented this in Simfactory3, and I definitively think this is the way to go. The current build recipes use Clang (not gcc, and not the system's "standard" compiler) for building. I did this because I was interested in C++11, and most system compilers (Intel, PGI, older versions of GCC, current Nvidia) don't support this. However, in the interest of backward compatibility we should probably also provide build recipes for (say) an older version of GCC that is supported by all compilers.
I will have to checkout and have a look at simfactory 3, so this may not be the best advised course of action ever devised. However: generally I would be happier if I had not to rely on simfactory (or any other complex tool) to build the external libraries. I tend to prefer my tools to do one thing and to that well. Simfactory (2) has become quite larger trying to do many things. I fear that building the external libraries such that they are compatible with the Cactus build (ie use the same OPENMP settings, debug and optimization settings) may pull in a lot of knowledge about the Cactus internal build system into simfactory which will make it complicated. I can understand Erik's desire to be able to build the ExternalLibraries independent of Cactus (for other projects) using the understanding of what it takes to build them gained while setting up the ExternalLibraries.
Also, not everyone is using simfactory. It may be possible to convince people to use it to build external libraries but I do not think that one can (nor should) try and force people to use simfactory to build Cactus and to submit jobs using it.
I am actually mostly happy with the current external libraries. My only desire would be for a collection of shell routines to be sourced to handle common tasks. This is somewhat similar to autoconf (ok everything there is eventually inlined into configure but at least on the m4 level there are libraries of functions) or the debian build system (which does source libraries of shell functions).
Basically I would suspect that a single mostly linear script is much easier to understand for newcomers than a system that uses fragments of scripts and a configuration file to stitch the fragments (and system supplied code) together.
I also agree that the later is quite likely "nicer" and better suited for large projects but and unsure if the dozen-or-so ExternalLibraries (each of which has its own quirks) qualify.
With respect to the non-standard options used to build Cactus: are those not also needed for the ExternalLibraries (if only to properly link with Cactus)?
Please note that I have not problem (at all) if simfactory 3 is used *by* Cactus as a library (with extra code added to get eg the compile options from the Cactus build system), my only concerns is using it to drive the Cactus build process (and the ExternalLibraries) thus having to know what Cactus will do.
Yours, Roland
On Mar 25, 2014, at 11:37 , Roland Haas rhaas@tapir.caltech.edu wrote:
Hello all,
I know this because I implemented this in Simfactory3, and I definitively think this is the way to go. The current build recipes use Clang (not gcc, and not the system's "standard" compiler) for building. I did this because I was interested in C++11, and most system compilers (Intel, PGI, older versions of GCC, current Nvidia) don't support this. However, in the interest of backward compatibility we should probably also provide build recipes for (say) an older version of GCC that is supported by all compilers.
I will have to checkout and have a look at simfactory 3, so this may not be the best advised course of action ever devised. However: generally I would be happier if I had not to rely on simfactory (or any other complex tool) to build the external libraries. I tend to prefer my tools to do one thing and to that well. Simfactory (2) has become quite larger trying to do many things. I fear that building the external libraries such that they are compatible with the Cactus build (ie use the same OPENMP settings, debug and optimization settings) may pull in a lot of knowledge about the Cactus internal build system into simfactory which will make it complicated. I can understand Erik's desire to be able to build the ExternalLibraries independent of Cactus (for other projects) using the understanding of what it takes to build them gained while setting up the ExternalLibraries.
No, there is nothing Cactus specific in Simfactory3's build system for external libraries.
I will be happy to show you how external libraries are built. In principle, it's just "download, unzip, configure, make, make test, make install, test", but in practice almost all packages need quite a bit more work.
You don't need to use Simfactory for anything else. This mechanism just installs external libraries, that's all it does.
However, there are a few interesting bits that Simfactory provides that are quite useful when building external libraries. You could specify these things manually when you build, or you could devise your own database holding these settings, if you want: - the directory where to install things - how the "make" command is called - whether (and which) modules to load before building
I am actually mostly happy with the current external libraries. My only desire would be for a collection of shell routines to be sourced to handle common tasks. This is somewhat similar to autoconf (ok everything there is eventually inlined into configure but at least on the m4 level there are libraries of functions) or the debian build system (which does source libraries of shell functions).
Shell routines are fine. However, I prefer Python, and have thus written a Python module for these common tasks.
Basically I would suspect that a single mostly linear script is much easier to understand for newcomers than a system that uses fragments of scripts and a configuration file to stitch the fragments (and system supplied code) together.
You would be surprised.
Let's talk about downloading only, ignoring patching/configuring/building/testing/installing etc. for the moment.
What command do you use to download? Let's assume you go for "wget", and let's assume it is available everywhere.
Now, you probably don't want to re-download a tarball that is already there. You can use if statements and tests for that, but luckily, wget has an option for this, "-c", so you use that.
Next you notice that wget doesn't always download tarballs when the url starts with https. Many system administrators can't seem to get their certificates up to date. So, you need to add --no-check-certificate. (Plus, you are not really interested in a secure connection to the web server, you really want end-to-end security, which only a GPG signature or checking md5sums will get you, to https doesn't help much anyway.)
Next you notice that Sourceforge (sic!) sometimes goes into endless redirection loops when the wget "-c" option is used. So you need to call wget twice: First with "-c"; then, if that fails, retry without "-c". That's inconvenient, but what can you do.
Next you find that wget chooses a bad file name when downloading *.zip files. So you need to use the "-O" option. That means you either need to specify the file name twice (as part of the url, and as argument to "-O"), or you need to use a function that automatically extracts the last part of the url. (Python provides such a function.)
Then, of course, there is the fact that some libraries (e.g. GCC) require multiple downloads. GCC requires six (!) different tarballs. So, open-coding all the wget magic gets cumbersome -- you want a list of urls, and a loop around the wget magic, or you need to define a shell function to wrap wget.
Now we can download tarballs. (Ian's Boost script that he posted earlier ago isn't even halfway there yet.) Next comes checking md5sums, unzipping, etc... None of these are as trivial as they should be. Things always break, and having something that "just works" requires quite a bit of non-trivial logic. A linear bash script isn't the way to go... Note that I always start out with linear bash scripts, and then always get cornered into writing something more complex.
I also agree that the later is quite likely "nicer" and better suited for large projects but and unsure if the dozen-or-so ExternalLibraries (each of which has its own quirks) qualify.
Give it a try. Pick an easy library (stay away from Boost, PETSc, HPX), and write a simple, linear script that works on all the systems that your group is using -- your local laptop and workstation, as well as the "usual" XSEDE and NERSC machines. (Leave out Crays and Blue Genes for now.) The script should download, unzip, configure, build, and install a library, and then run a simple test ensuring that the library has been installed fine, such as e.g. compiling and running a five-line C program that outputs the library's version number to stdout.
With respect to the non-standard options used to build Cactus: are those not also needed for the ExternalLibraries (if only to properly link with Cactus)?
You don't need these. A standard install will work fine, if you make sure that some basic things match (e.g. libstdc++ version).
Please note that I have not problem (at all) if simfactory 3 is used *by* Cactus as a library (with extra code added to get eg the compile options from the Cactus build system), my only concerns is using it to drive the Cactus build process (and the ExternalLibraries) thus having to know what Cactus will do.
It doesn't. I'm using Simfactory3 for non-Cactus projects as well. That works just fine. That's one of the reasons for Simfactory3. It just gives you some nice Python modules (sorry, no Bash scripts to pull in) to do simple tasks. Such as, "tell me what the make command is called on this system".
-erik
Hello Erik, all,
Shell routines are fine. However, I prefer Python, and have thus written a Python module for these common tasks.
Fine with me. I don't have a particular preference for the language used (as long as it is somewhat well known and reasonably suitable). Shell scripts seemed natural since lots of command calls and output parsing seems involved.
You would be surprised.
That seems quite likely given the list of issues one has to be aware of. I am rather naive in this respect. I'll have a look at simfactory so that at least I can provide constructive criticism :-).
Yours, Roland
While something like a SimFactory 3 based system might be the best in the long run, I think Ian's proposed download.sh system would be nice incremental improvement on the existing ExternalLibraries system. It allows us to have minimal changes to the existing mostly-working system, but with the added benefit of not having large files around when not necessary. For that reason alone, I think it would be good to implement it in the short to mid term, until something better is ready.
On 25 Mar 2014, at 15:58, Erik Schnetter schnetter@gmail.com wrote:
On Mar 25, 2014, at 10:30 , Ian Hinder ian.hinder@aei.mpg.de wrote:
On 25 Mar 2014, at 14:53, Frank Loeffler knarf@cct.lsu.edu wrote:
On Mon, Mar 24, 2014 at 04:07:35PM -0700, Roland Haas wrote:
According to the minutes from 2013-12-16 a decision was reached:
Thanks for digging this out. This is what I remembered.
What is not in the notes is that Ian also desired a way to force it to download so that for cluster where no network connection (or git) is available at build time (or preparing for a flight) one can force download on ones laptop then use simfactory to sync the tarballs.
I think we mentioned at least in the call that such a possibility should exist - an option to download all possibly required tarballs.
The obvious place to store the URL for the tarball is in the configuration.ccl file. What is the recommended procedure for reading entries from this file outside of the build system? It would have been nice to find that the Perl files in lib/sbin were Perl modules which could be used outside the build system, but I was disappointed to find this was not the case. The file is parsed by lib/sbin/ConfigurationParser.pl.
I think we want to move away from building external libraries as part of the Cactus build; instead, we want to build them ahead of time, e.g. via Simfactory. This has several advantages, such as e.g. that a "make clean" doesn't require rebuilding the libraries, and that several Cactus configurations can use the same external libraries, and that one can even have one power user build the libraries on a system while all others simply use them.
One additional advantage is that building external libraries from within Cactus is actually quite difficult, since the Cactus compiler options are often "strange" and need to be "cleaned" before one can use them. The respective shell code is arcane, and often breaks between different system. Building external libraries on their own, independent of Cactus, is much simpler.
I know this because I implemented this in Simfactory3, and I definitively think this is the way to go. The current build recipes use Clang (not gcc, and not the system's "standard" compiler) for building. I did this because I was interested in C++11, and most system compilers (Intel, PGI, older versions of GCC, current Nvidia) don't support this. However, in the interest of backward compatibility we should probably also provide build recipes for (say) an older version of GCC that is supported by all compilers.
I thought about that this morning, and agree that we need a better system for the external libraries. I don't know if this should be tied to simfactory, or should be embedded in Cactus.
However, in the short term, before simfactory 3 is deployed (and works with gcc), I would like to patch the current Cactus mechanism. A similar idea might be useful for initial data thorns which come with large binary datasets; the datasets might be better distributed only to the machines which need them, rather than making a detour via my home internet connection.
I have filtered the Cactus Boost git repository to remove the tarball, and added code to perform the download before building if needed. The tarball is currently hosted in the downloads section of the bitbucket repository, but this could change. Consider this a proof of concept.
https://bitbucket.org/ianhinder/boost
This thorn puts the tarball in Cactus/libcache, and reuses it if it find it there already. It checks the md5sum of the downloaded file against a value hard-coded in the script. It deletes old versions of the tarball that it finds before moving the new version into place. It has been tested on Mac OS 10.8.3 (my laptop) and Scientific Linux 6.0 (Datura). The download code is currently in the thorn, but should move into Cactus once it is shared between thorns. You can initiate the download for this thorn by running the script, so you could do, for example,
for x in arrangements/*/*/download.sh; do $x; done
We could provide a make target or script to do this.
At the moment, the download script needs several thorn-specific pieces of information which are currently hard-coded:
DISTBASE=boost DISTVERSION=1_55_0 DISTNAME=${DISTBASE}_${DISTVERSION}.tar.gz DISTHASH=93780777cfbf999a600f62883bd54b17 DISTSRC=https://bitbucket.org/ianhinder/boost/downloads
I think these should go into the configuration.ccl file of the thorn, and the download script can then be a part of Cactus, reading the required information for each thorn. However, I don't know the correct way to read this information from the CCL file from a shell script. I could knock up something which worked, but which probably wouldn't be very robust. It would be better for Cactus to provide an API for getting this sort of information.
users@lists.einsteintoolkit.org