Hi
Please consider joining the weekly Einstein Toolkit phone call at 10 am US central time on Mondays. As usual, you can find instructions how to join on the following web site:
http://einsteintoolkit.org/community/support/
In short: the number is (+1) 225-578-4942 or (+1) 866-573-0359 and the conference id is 118682#.
In addition to anything that might be brought up at the meeting, the primary issue at the moment is to get the code ready for a new release in May, with the planned schedule shown below. At the moment this means to test, test and test. Bugs and smallish new features might still go in, but larger additions (like new thorns) are already 'out'.
Last week a number of maintainers agreed to test machines. So far, only four are in:
https://docs.einsteintoolkit.org/et-docs/Release_Details
We currently miss:
Ian: datura Erik: bluewaters, hopper, orca, surveyor, vesta Roland: kraken, zwicky Peter: pandora, queenbee, tezpur Frank: philip (speed problems) Tanja: stampede Bruno: tresles
plus a couple of private machines. Please update the results, or the wiki indicating if you cannot do the tests on a particular machine.
The currently very scarce results indicate, that at least on one machine (spine - a Debian workstation, using gcc 4.7) all tests pass. On others quite some number fail (e.g. supermuc). Common problems seem to be one Exact testsuite, and some GRHydro suites, especially the new WENO test stuite. Symmetry thorns also fail some of their tests on more than one of the four machines.
As always: failure don't necessarily mean that something is wrong. Sometimes, the tolerances are just too strict for the underlying code, and CPUs, different compilers or compiler flags then show these differences.
So, in short: the good news is that at least one machine doesn't show test suite failures, something that was often not the case in the past. Now we 'only' need to figure out what needs to be done with tests that fail on some of the machines - especially the large production clusters.
We also mentioned possibilities for the new release name last time. The current list is on the wiki as well. Please add other names if you like:
https://docs.einsteintoolkit.org/et-docs/Release_Details
- ET_2013_05 schedule - Apr 08: No new major additions to the ET until unfreeze - May 01: feature freeze of /trunk, extensive testing begins (test-suites on all relevant machines) - May 10: release branches are created trunk unfrozen again release-branch code in only-important-bugs-freeze bug-fixing primarily in release branch - May 21: complete freeze, to prepare tar-balls, add tags ect. - May 23: target release date (or earlier if all bugs are fixed)
Frank Löffler
On Sun, Apr 28, 2013 at 11:44 PM, Frank Loeffler knarf@cct.lsu.edu wrote:
The currently very scarce results indicate, that at least on one machine
(spine - a Debian workstation, using gcc 4.7) all tests pass. On others quite some number fail (e.g. supermuc). Common problems seem to be one Exact testsuite, and some GRHydro suites, especially the new WENO test stuite. Symmetry thorns also fail some of their tests on more than one of the four machines.
As always: failure don't necessarily mean that something is wrong. Sometimes, the tolerances are just too strict for the underlying code, and CPUs, different compilers or compiler flags then show these differences.
So, in short: the good news is that at least one machine doesn't show test suite failures, something that was often not the case in the past.
This is not really news, given that all test cases are run automatically after every commit: https://trac.einsteintoolkit.org/; look for "Status". In the past months, we had more problems with repositories (and the automated mechanism itself) than with the Einstein Toolkit code base.
-erik
On Sun, Apr 28, 2013 at 11:58:00PM -0400, Erik Schnetter wrote:
This is not really news, given that all test cases are run automatically after every commit: https://trac.einsteintoolkit.org/; look for "Status".
You are right. This has been steady for some months now, in no small part due to these automated tests - thanks to Barry and Ian. I was referring to a more distant past, when it was a bit like this:
http://geek-and-poke.com/2011/08/hudson-status-monitor.html
Speaking about tests: I did a quick test on spine (the machine that currently doesn't show failures). This uses gcc, and I only changed the optimization from -O2 to -Ofast. The result is some testsuite failures, most of them the 'usual set': Exact, GRHydro: weno, Symmetry thorns.
-Ofast (apart other things) uses some non-conforming math, something that is not default for gcc. The intel compiler, on the other hand, does some non-conforming math by default (but very likely not the same subset as gcc with -Ofast). Part of the testsuite failures could just be due to this: different compiler defaults and/or options. Also here: this is not news. We are likely not talking about a lot of bugs in the code, but more likely about how much inaccuracy we are willing to trade for speed.
Frank
On 29 Apr 2013, at 06:55, Frank Loeffler knarf@cct.lsu.edu wrote:
On Sun, Apr 28, 2013 at 11:58:00PM -0400, Erik Schnetter wrote:
This is not really news, given that all test cases are run automatically after every commit: https://trac.einsteintoolkit.org/; look for "Status".
You are right. This has been steady for some months now, in no small part due to these automated tests - thanks to Barry and Ian. I was referring to a more distant past, when it was a bit like this:
Without automated testing, it used to be the case that the full set of tests would likely not be run between two releases. So one had to fix all the problems at once, and there was no indication of which commit caused the failure. Over the past year, as we have brought the automated testing system online, we have usually flagged these issues instantly, and it is usually very clear which commit has broken the tests. I'm not sure I agree with Erik's statement that we have had more problems with the repositories and automated test system than with the code. People have frequently committed code which broke some of the tests (which usually results in a private email from me); it's just that the code problems have been fixed very quickly after being discovered as it was clear what caused them. Ideally, the tests would be easy enough to run that everyone would run them before committing, but this is not the case at the moment, as they take too long.
If we consider the automated test system stable enough now, we can switch to a mode where the test failures are reported to a mailing list (maybe the commits mailing list, or a new one) and the original committer, so that everyone is aware of the problems who wants to be.
On Apr 29, 2013, at 5:44 AM, Frank Loeffler wrote:
The currently very scarce results indicate, that at least on one machine (spine - a Debian workstation, using gcc 4.7) all tests pass. On others quite some number fail (e.g. supermuc).
On supermuc, the failures on 1 process are really due to the fact that I haven't figured how to properly run the testsuites with the simfactory mechanism. I tried to tweak the instructions at:
https://docs.einsteintoolkit.org/et-docs/Testsuite_Machines
which are obviously outdated, but to no avail so far. The tests are always run with two processes, so I've removed the supermuc__1_8.log file for now.
Perhaps this discussion could be added to today's list of topics.
We also mentioned possibilities for the new release name last time. The current list is on the wiki as well. Please add other names if you like:
I've added the name Planck, mostly to celebrate the mission results that came out last month:
Best, Eloisa
On 29 Apr 2013, at 05:44, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi
Please consider joining the weekly Einstein Toolkit phone call at 10 am US central time on Mondays. As usual, you can find instructions how to join on the following web site:
http://einsteintoolkit.org/community/support/
In short: the number is (+1) 225-578-4942 or (+1) 866-573-0359 and the conference id is 118682#.
In addition to anything that might be brought up at the meeting, the primary issue at the moment is to get the code ready for a new release in May, with the planned schedule shown below. At the moment this means to test, test and test. Bugs and smallish new features might still go in, but larger additions (like new thorns) are already 'out'.
I've updated the instructions for running the tests to use simfactory 2 and its support for running tests:
https://docs.einsteintoolkit.org/et-docs/Testsuite_Machines
You can add machine-specific notes to that page, but there shouldn't be any. I haven't run through and tested the instructions; if there is something wrong, please fix it or let me know.
On Mon, Apr 29, 2013 at 11:28 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 29 Apr 2013, at 05:44, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi
Please consider joining the weekly Einstein Toolkit phone call at 10 am US central time on Mondays. As usual, you can find instructions how to join on the following web site:
http://einsteintoolkit.org/community/support/
In short: the number is (+1) 225-578-4942 or (+1) 866-573-0359 and the conference id is 118682#.
In addition to anything that might be brought up at the meeting, the primary issue at the moment is to get the code ready for a new release in May, with the planned schedule shown below. At the moment this means to test, test and test. Bugs and smallish new features might still go in, but larger additions (like new thorns) are already 'out'.
I've updated the instructions for running the tests to use simfactory 2 and its support for running tests:
https://docs.einsteintoolkit.org/et-docs/Testsuite_MachinesYou can add machine-specific notes to that page, but there shouldn't be any. I haven't run through and tested the instructions; if there is something wrong, please fix it or let me know.
If, for some reason, Simfactory doesn't work on any of the tested (public) machines, then please make a very prominent note and open a ticket. We collect instructions for using the machines in Simfactory, and anybody with access to a machine where we claim the ET has been tested will be expecting that things work there out of the box.
-erik
users@lists.einsteintoolkit.org