ExternalLibraries thorns containing tarballs can be very large, in particular if the history of the thorn grows and contains multiple tarballs.
I have just tested storing the source tree instead of the tarball in ExternalLibraries/Boost. This seems to work surprisingly well. For example, boost_1_54_0.tar.gz (the current version) has 66 MByte, while a git repository containing the last five versions of Boost as uncompressed source trees has 75 MByte. Clearly, this is very good.
Other advantages of storing source trees are: - this is what we do for every thorn, so why not for ExternalLibraries (d'oh) - one can easily look at the source code of an external library - no need to untar and later delete the source tree while building - no need to apply patches -- the changes can be made directly in the source tree
In my opinion, this is the way to go. I am currently setting up a new repository of ExternalLibraries/Boost in this way. (Obviously, it is not possible to keep the current repository, since its history is already at 406 MB.)
-erik
Hi,
On Thu, Aug 08, 2013 at 08:34:09PM -0400, Erik Schnetter wrote:
ExternalLibraries thorns containing tarballs can be very large, in particular if the history of the thorn grows and contains multiple tarballs.
It only really grows (for the user) if a revision control system is used that stores all of the history on the client side - something a user wouldn't be interested in usually, for external libraries. git is one of these. A counter example is Subversion, which wouldn't have that problem, and all external libraries are currently stored in Subversion.
Other advantages of storing source trees are:
- this is what we do for every thorn, so why not for ExternalLibraries (d'oh)
External libraries are not like other thorns. Even when present in a thornlist, the source of external libraries might not be used in the end if a system versions is detected and usable, but it needs to be present. Decompressing all external libraries would lead to even more small files, which especially on clusters could get you into trouble. Of course you need to have the inodes when you compile it anyway, but again: you might not need to (compile it), and most of these are deleted again after the build.
- one can easily look at the source code of an external library
- no need to untar and later delete the source tree while building
For a regular user this happens automagically, but I agree that this would be a plus for the maintainer of that thorn.
- no need to apply patches -- the changes can be made directly in the source tree
I don't think this would be good. I would like to see which changes the Cactus thorn did compared to the vanilla version. Also, from a developers standpoint maintaining patches is far easier for upgrading the library than recreating and applying patches every time for an upgrade by hand. I strongly urge everyone to keep the vanilla version clearly separated from Cactus-specific changes.
In my opinion, this is the way to go. I am currently setting up a new repository of ExternalLibraries/Boost in this way. (Obviously, it is not possible to keep the current repository, since its history is already at 406 MB.)
I don't really see a problem with the current setup. Even the boost svn checkout shouldn't be that large. The problem you describe really only applies for a git repository. We don't need to use git. We even don't use git for the external libraries.
Frank
On 9 Aug 2013, at 02:46, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi,
On Thu, Aug 08, 2013 at 08:34:09PM -0400, Erik Schnetter wrote:
ExternalLibraries thorns containing tarballs can be very large, in particular if the history of the thorn grows and contains multiple tarballs.
It only really grows (for the user) if a revision control system is used that stores all of the history on the client side - something a user wouldn't be interested in usually, for external libraries. git is one of these. A counter example is Subversion, which wouldn't have that problem, and all external libraries are currently stored in Subversion.
Taking a set of source trees which are very similar and compressing each of them into a tarball makes it impossible to efficiently delta-compress them. This is what SVN does in its repository. While most users probably don't care about what happens hidden inside the server-hosted SVN repository, the mere fact that there is this duplication suggests that it is an inelegant, and hence probably wrong, solution to the problem. Note that users will have to check out the whole new version of the library each time it changes, rather than just downloading the differences. This also applies to syncing to remote machines. So users will see a performance hit with the current SVN approach anyway.
Other advantages of storing source trees are:
- this is what we do for every thorn, so why not for ExternalLibraries (d'oh)
External libraries are not like other thorns. Even when present in a thornlist, the source of external libraries might not be used in the end if a system versions is detected and usable, but it needs to be present. Decompressing all external libraries would lead to even more small files, which especially on clusters could get you into trouble. Of course you need to have the inodes when you compile it anyway, but again: you might not need to (compile it), and most of these are deleted again after the build.
This is a valid concern. Erik, with your testing of boost, could you measure the number of files in the working tree, and the number of files in the built tree, for just the boost library? I imagine that the number of files in the built tree is a factor of a few larger than the number in the source tree. However, as Frank said, if the user is not building this library, then they do have a much larger inode cost if we store the source tree. We discusses skipping syncing of certain external libraries before; maybe that is a better solution here.
Note that I am already opposed to using a large library such as Boost in the ET. It requires a GB of space to uncompress, and more to build, for little benefit that I can see (though I have not looked). If Boost is licensed appropriately and written in a modular-enough fashion, maybe it is possible to extract just the bits that people find useful? I'm sure we are pulling in a huge amount of code that we don't need by including Boost.
- one can easily look at the source code of an external library
- no need to untar and later delete the source tree while building
For a regular user this happens automagically, but I agree that this would be a plus for the maintainer of that thorn.
- no need to apply patches -- the changes can be made directly in the source tree
I don't think this would be good. I would like to see which changes the Cactus thorn did compared to the vanilla version.
I assume Erik planned to keep the original version on a branch, and have an ET branch with our changes.
Also, from a developers standpoint maintaining patches is far easier for upgrading the library than recreating and applying patches every time for an upgrade by hand. I strongly urge everyone to keep the vanilla version clearly separated from Cactus-specific changes.
A developer would not have to work with patches. You would commit your changes on the ET branch of the library, and use the version control tool to see differences, merge in the new version, etc. This is much better than messing about with patches, which are just a hack used in the absence of proper version control.
In my opinion, this is the way to go. I am currently setting up a new repository of ExternalLibraries/Boost in this way. (Obviously, it is not possible to keep the current repository, since its history is already at 406 MB.)
I don't really see a problem with the current setup. Even the boost svn checkout shouldn't be that large. The problem you describe really only applies for a git repository. We don't need to use git. We even don't use git for the external libraries.
There are problems with the current approach when using SVN:
* Updating an external library requires the whole library to be downloaded, rather than just what has changed * Syncing to a remote cluster requires the whole library to be sent, rather than what has changed (often on a residential internet connection with limited upstream bandwidth)
Most of the active ET developers prefer to work with git anyway, and since git stores all the history locally, the problem is worse for us. Note that the boost external library thorn is in git, which is why Erik noticed this problem in the first place.
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
Hello all,
Taking a set of source trees which are very similar and compressing each of them into a tarball makes it impossible to efficiently delta-compress them.
This is not quite true. gzip (at least version 1.6 on Debian) has an option rsyncable which claims to help in these situations, though most likely not as much as actually checking in the source code into the repository.
This is a valid concern. Erik, with your testing of boost, could you measure the number of files in the working tree, and the number of files in the built tree, for just the boost library? I imagine that the number of files in the built tree is a factor of a few larger than the number in the source tree. However, as Frank said, if the user is not building this library, then they do have a much larger inode cost if we store the source tree. We discusses skipping syncing of certain external libraries before; maybe that is a better solution here.
I would like to also lend my voice to this argument. Having the tarballs extracted in the checkout creates very many small files which can be terribly slow to synchronize via rsync. Much slower than transferring a single medium sized (~30MB is the largest tarball, excluding boost) file.
Note that I am already opposed to using a large library such as Boost in the ET. It requires a GB of space to uncompress, and more to build, for little benefit that I can see (though I have not looked). If Boost is licensed appropriately and written in a modular-enough fashion, maybe it is possible to extract just the bits that people find useful? I'm sure we are pulling in a huge amount of code that we don't need by including Boost.
I thin the Boost ExternalLibrary in its current repository is not a good example. We would not handle it this way. Git is a not a good VCS for ExternalLibraries since it keeps the whole history around which we don't need for the external libraries. Boost *can* be split, eg Debian delivers it as ~13 packages. Not sure though if it is worthwhile for us to split it up like this. I'd much rather have the checkout commented out in the thornlist (same as OpenCL) and then if someone needs it they can enable it. If a thorn then relies on Boost the thorn can include the corresponding source files if that is sufficient (and once sufficiently many thorns do so we might as well just include Boost itself...).
A developer would not have to work with patches. You would commit your changes on the ET branch of the library, and use the version control tool to see differences, merge in the new version, etc. This is much better than messing about with patches, which are just a hack used in the absence of proper version control.
I would expect we can already do so with the current setup when eg the maintainer of the ExternalLibrary keeps a (local) git repository (eg via git-svn) and rebase the ET specific patches after each update. git-format-patch then produces the patches that need applying. This is more or less how I handle eg. my own Carpet branch where I have local modifications that I like but which were rejected from upstream so I cannot push them.
Most of the active ET developers prefer to work with git anyway, and since git stores all the history locally, the problem is worse for us. Note that the boost external library thorn is in git, which is why Erik noticed this problem in the first place.
I believe git is not a good choice for ExternalLibraries. In my understanding we do not want to modify ExternalLibraries's source code at all if possible since otherwise we run the risk of not able to function with at plain vanille copy already installed on the system. If we want to provide extra functionality beyond what the library offers, I'd rather use a second thorn to do so. From my point of view ExternalLibraries are just packaged up external components to be used with the ET, not something that the ET develops. Eg. if I was actually developing for LORENE, I would not use the ET tarball but download it on my own.
Yours, Roland
- -- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
Hi,
On Sun, Aug 11, 2013 at 04:59:47PM +0200, Ian Hinder wrote:
Taking a set of source trees which are very similar and compressing each of them into a tarball makes it impossible to efficiently delta-compress them.
That's not completely true. 'tar' isn't really the problem. The compression we use is.
This is what SVN does in its repository. While most users probably don't care about what happens hidden inside the server-hosted SVN repository, the mere fact that there is this duplication suggests that it is an inelegant, and hence probably wrong, solution to the problem.
We knew that it is not the most elegant solution from the beginning. However, it is the most convenient for a user.
Note that users will have to check out the whole new version of the library each time it changes, rather than just downloading the differences. This also applies to syncing to remote machines. So users will see a performance hit with the current SVN approach anyway.
Yes. But then we argued that the external libraries aren't going to change so often anyway - and when you actually have an update, you are probably connected to a high-speed network at work anyway.
This is a valid concern. Erik, with your testing of boost, could you measure the number of files in the working tree, and the number of files in the built tree, for just the boost library? I imagine that the number of files in the built tree is a factor of a few larger than the number in the source tree. However, as Frank said, if the user is not building this library, then they do have a much larger inode cost if we store the source tree. We discusses skipping syncing of certain external libraries before; maybe that is a better solution here.
Also, the files in the build tree are deleted once the library is built, leading to much lower overhead especially with a lot of configurations that might need different built versions of the libraries.
Note that I am already opposed to using a large library such as Boost in the ET. It requires a GB of space to uncompress, and more to build, for little benefit that I can see (though I have not looked).
I didn't look as well, but hear almost every time how bad of a decision some people think it was to have certain other projects relying so heavily on boost.
I assume Erik planned to keep the original version on a branch, and have an ET branch with our changes.
That would limit the problem somewhat. But still - as user I might be interested to see which changes are necessary to make library X work with the ET without digging into the documentation of a VCS to find the patches with their descriptions - especially if they evolved over time and I am not at all interested in that evolution, just the actual patch, and reason.
A developer would not have to work with patches. You would commit your changes on the ET branch of the library, and use the version control tool to see differences, merge in the new version, etc. This is much better than messing about with patches, which are just a hack used in the absence of proper version control.
When tracking an externally developed software I would always like to minimize the changes I have - dropping patches with time if they are not necessary anymore. I would like to keep patches apart that have distinct purposes. I am usually not very much interested in how these patches evolve in time. All I really care about is a minimal set of patches with a distinct, stated, reason for the each upstream version.
Correct me if I am wrong, but what you describe is something different, isn't it?
- Updating an external library requires the whole library to be downloaded, rather than just what has changed
- Syncing to a remote cluster requires the whole library to be sent, rather than what has changed (often on a residential internet connection with limited upstream bandwidth)
These two are really the same, single issue. Also - don't sync when you are on low bandwidth. The same would apply for an initial clone of a giant git repository, including history you don't care about.
Frank
users@lists.einsteintoolkit.org