Hi
Please consider joining the weekly Einstein Toolkit phone call at
10 am US central time on Mondays. As usual, you can find instructions
how to join on the following web site:
http://einsteintoolkit.org/community/support/
I short: the number is (+1) 225-578-4942 or (+1) 866-573-0359 and
the conference id is 118682#.
My agenda contains:
- Fall workshop
- PLEASE let us know if you come [1]!
- Hotel booking
- Release planning
- release date: Mon Oct 24 2011
- freeze date: Mon Sep 26 - the day of the call!
- branch name: ET_2011_10
- List of release-concerning bugs [2]
- ET paper: proposed "freeze" by Oct 3rd (Mon after this call)
As always: feel free to add to this list.
Frank Loeffler
[1] https://docs.einsteintoolkit.org/et-docs/ET_Workshop_Fall_2011#Fall_Einstei…
[2] https://trac.einsteintoolkit.org/query?status=!closed&milestone=ET_2011_11&…
Hi,
I did some strong scaling tests on Stampede (https://www.xsede.org/tacc-stampede) with both Whisky and GRHydro (using the development version of ET in both cases). I used a Carpet par file that Roberto DePietri provided me and that he used for similar tests on an Italian cluster (I have attached the GRHydro version, the Whisky one is similar except for using Whisky instead of GRHydro). I used both Intel MPI and Mvapich and I did both pure MPI and MPI/OpenMP runs.
I have attached a text file with my results. The first column is the name of the run (if it starts with mvapich it used mvapich otherwise it used Intel MPI), the second one is the number of cores (option --procs in simfactory), the third one the number of threads (--num-threads), the fourth one the time in seconds spent in CCTK_EVOL, and the fifth one the walltime in seconds (i.e., the total time used by the run as measured on the cluster). I have also attached a couple of figures that show CCTK_EVOL vs #cores and walltime vs #cores (only for Intel MPI runs).
First of all, in pure MPI runs (--num-threads=1) I was unable to run on more than 1024 cores using Intel MPI (the run was just crashing before iteration zero or hanging up). No problem instead when using --num-threads=8 or --num-threads=16. I also noticed that scaling was particularly bad in pure MPI runs and that a lot of time was spent outside CCTK_EVOL (both with Intel MPI and MVAPICH). After speaking with Roberto, I found out that the problem is due to 1D ASCII output (which is active in that parfile) and that makes the runs particularly slow above ~100 cores on this machine. In plot_scaling_walltime_all.pdf I plot also two pure MPI runs, but without 1D ASCII output and the scaling is much better in this case (the time spent in CCTK_EVOL is identical to the case with 1D output and hence I didn't plot them in the other figure). I didn't try using 1D hdf5 output instead, does anyone use it?
According to my tests, --num-threads=16 performs better than --num-threads=8 (which is the current default value in simfactory) and Intel MPI seems to be better than MVAPICH. Is there a particular reason for using 8 instead of 16 threads as the default simfactory value on Stampede?
Let me know if you have any comment or suggestion.
Cheers,
Bruno
Dr. Bruno Giacomazzo
JILA - University of Colorado
440 UCB
Boulder, CO 80309
USA
Tel. : +1 303-492-5170
Fax : +1 303-492-5235
email : bruno.giacomazzo(a)jila.colorado.edu
web: http://www.brunogiacomazzo.org
----------------------------------------------------------------------
There are only 10 types of people in the world:
Those who understand binary, and those who don't
----------------------------------------------------------------------
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1
Present: Ian, Eloisa, Seth, Josh, Yosef, Erik, Frank, Peter,
Roland, Steve, Matt Kinsey
ET release:
* most tests pass on all machines, exceptions are GRHydro's weno and
slow_sector tests. Roland will look into them (since he put them into
the repo). Suspicions are that they were generated using gcc but are
now tested using intel (thing --fast-math)
* Peter reports some test failures with new intel compilers as well
(on his laptop)
* currently have 34 open tickets flagged for the release. Please have
a look. 6 of them are in review, some are even reviewed ok.
* Ian now automatically builds documentation for ET, will not use it
on ET page yet since it is build from trunk and the page should be the
release code. Ian will also build the stable documentation and we can
link to that one.
* Frank to ask Bruno to give the documentation a quick check
Yours,
Roland
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)
Comment: Using GnuPG with Mozilla - http://enigmail.mozdev.org/
iEYEARECAAYFAlF+l+cACgkQTiFSTN7SboXvLQCgzQIShmcRP6XstLMRuo4sP7Yd
HX4AnjbevIB8nzB8tyivyH/suZuKN5O0
=1Ss0
-----END PGP SIGNATURE-----
The Llama Multi-Block Infrastructure for Cactus <http://llamacode.org/> is
now publicly available under the GNU General Public License. Llama
provides three-dimensional multi-block capability for Cactus-based
simulations that can be combined with Carpet's adaptive mesh refinement
functionality. Llama decomposes the domain into multiple (potentially
overlapping) blocks with different local coordinate systems. This allows
e.g. spherical domains, spherical excision, adaptive radial/angular
resolution, etc., without incurring coordinate singularities.
Llama provides several patch systems suitable for single and binary objects
in relativistic astrophysics, and is well integrated with the Einstein
Toolkit <http://einsteintoolkit.org/>. Llama was already used for several
publications <http://llamacode.org/research.html>, and we believe the code
is ready to be used in other projects. We are seeking volunteers to help us
add tutorials and documentation, improve error messages, and generally
shake down and brush up the code for a future inclusion in the Einstein
Toolkit.
To aid others in getting started using Llama, we will be hosting a virtual
workshop where we provide an overview of the code and answer questions.
Details will be announced shortly.
Llama constitutes the fruit of a significant effort of several people over
several years. We make Llama public to help modernize the computational
tools used in our community, and in the hope to boost Llama itself by
inviting contributions from everybody. We ask you to acknowledge our effort
by following the citation guidelines described on <
http://llamacode.org/research.html>.
The Llama groomers:
R. Haas, I. Hinder, D. Pollney, C. Reisswig, E. Schnetter, B. Wardell
--
Erik Schnetter <schnetter(a)cct.lsu.edu>
http://www.perimeterinstitute.ca/personal/eschnetter/
Hi
Please consider joining the weekly Einstein Toolkit phone call at
10 am US central time on Mondays. As usual, you can find instructions
how to join on the following web site:
http://einsteintoolkit.org/community/support/
In short: the number is (+1) 225-578-4942 or (+1) 866-573-0359 and
the conference id is 118682#.
In addition to anything that might be brought up at the meeting, the
primary issue at the moment is to get the code ready for a new release
in May, with the planned schedule shown below. At the moment this means
to test, test and test. Bugs and smallish new features might still go
in, but larger additions (like new thorns) are already 'out'.
Last week a number of maintainers agreed to test machines. So far, only
four are in:
https://docs.einsteintoolkit.org/et-docs/Release_Details
We currently miss:
Ian: datura
Erik: bluewaters, hopper, orca, surveyor, vesta
Roland: kraken, zwicky
Peter: pandora, queenbee, tezpur
Frank: philip (speed problems)
Tanja: stampede
Bruno: tresles
plus a couple of private machines. Please update the results, or the
wiki indicating if you cannot do the tests on a particular machine.
The currently very scarce results indicate, that at least on one machine
(spine - a Debian workstation, using gcc 4.7) all tests pass. On others
quite some number fail (e.g. supermuc). Common problems seem to be one
Exact testsuite, and some GRHydro suites, especially the new WENO test
stuite. Symmetry thorns also fail some of their tests on more than one
of the four machines.
As always: failure don't necessarily mean that something is wrong.
Sometimes, the tolerances are just too strict for the underlying code,
and CPUs, different compilers or compiler flags then show these
differences.
So, in short: the good news is that at least one machine doesn't show
test suite failures, something that was often not the case in the past.
Now we 'only' need to figure out what needs to be done with tests that
fail on some of the machines - especially the large production clusters.
We also mentioned possibilities for the new release name last time. The
current list is on the wiki as well. Please add other names if you like:
https://docs.einsteintoolkit.org/et-docs/Release_Details
- ET_2013_05 schedule
- Apr 08: No new major additions to the ET until unfreeze
- May 01: feature freeze of /trunk, extensive testing begins
(test-suites on all relevant machines)
- May 10: release branches are created trunk unfrozen again
release-branch code in only-important-bugs-freeze bug-fixing
primarily in release branch
- May 21: complete freeze, to prepare tar-balls, add tags ect.
- May 23: target release date (or earlier if all bugs are fixed)
Frank Löffler
#590: McLachlan should allow other thorns to set the gauge
------------------------------------+---------------------------------------
Reporter: bmundim | Owner: diener
Type: defect | Status: assigned
Priority: major | Milestone: ET_2013_05
Component: EinsteinToolkit thorn | Version:
Resolution: | Keywords:
------------------------------------+---------------------------------------
Comment (by eschnett):
This is a fairly invasive change to a very fundamental thorn. Such a
change needs thorough testing.
Peter says above that he didn't run the full set of test cases yet.
This patch is also missing new test cases, both for the general case
(everything handled by McLachlan) as well as the new cases (either lapse
and/or shift handled by another mechanism). Presumably, one would not just
record the new results, but also test that the results are correct, e.g.
ensuring convergence with a gauge wave test case.
What other thorn would make use of this functionality? I assume that those
who asked for this functionality would be happy to test it with their
additional gauge thorns (even if those are private thorns). If there is no
one volunteering for such testing, then it would seem that we are
developing functionality that is not actually currently needed.
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/590#comment:19>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
#590: McLachlan should allow other thorns to set the gauge
------------------------------------+---------------------------------------
Reporter: bmundim | Owner: diener
Type: defect | Status: assigned
Priority: major | Milestone: ET_2013_05
Component: EinsteinToolkit thorn | Version:
Resolution: | Keywords:
------------------------------------+---------------------------------------
Comment (by knarf):
Peter, can you please comment? Are there any reasons not to include this
so far?
--
Ticket URL: <https://trac.einsteintoolkit.org/ticket/590#comment:18>
Einstein Toolkit <http://einsteintoolkit.org>
The Einstein Toolkit
A general note:
We do not want to run the auto-detection when *_DIR is NO_BUILD. To run the
auto-detection, leave this variable blank. NO_BUILD is reserved for cases
where the library is known to be found by means opaque to the configure
script (e.g. MPI that is always present on Crays).
There was a discussion about this in Trac, and we initially and erroneously
agreed that running the auto-detection would be a good idea. It turns out
it wasn't (it broke the build on several systems), and thus we reverted
this change.
-erik
On Mon, Apr 22, 2013 at 10:46 AM, <knarf(a)cct.lsu.edu> wrote:
> User: knarf
> Date: 2013/04/22 09:46 AM
>
> Modified:
> /trunk/
> configure.sh
>
> Log:
> make configure look harder for installed fftw3
>
> File Changes:
>
> Directory: /trunk/
> ==================
>
> File [modified]: configure.sh
> Delta lines: +31 -10
> ===================================================================
> --- trunk/configure.sh 2013-04-10 15:13:01 UTC (rev 21)
> +++ trunk/configure.sh 2013-04-22 14:46:28 UTC (rev 22)
> @@ -11,31 +11,52 @@
> set -e # Abort on errors
>
>
> -
>
> ################################################################################
> # Search
>
> ################################################################################
>
> -if [ -z "${FFTW3_DIR}" ]; then
> +if [ -z "${FFTW3_DIR}" \
> + -o "$(echo "${FFTW3_DIR}" | tr '[a-z]' '[A-Z]')" = 'NO_BUILD' ]
> +then
> echo "BEGIN MESSAGE"
> echo "FFTW3 selected, but FFTW3_DIR not set. Checking some places..."
> echo "END MESSAGE"
> -
> - FILES="include/fftw3.h"
> - DIRS="/usr /usr/local /usr/local/fftw3 /usr/local/packages/fftw3
> /usr/local/apps/fftw3 ${HOME} c:/packages/fftw3"
> +
> + DIRS="/usr /usr/local /usr/local/packages /usr/local/apps /opt/local
> ${HOME} c:/packages"
> for dir in $DIRS; do
> - FFTW3_DIR="$dir"
> - for file in $FILES; do
> - if [ ! -r "$dir/$file" ]; then
> - unset FFTW3_DIR
> + DIRS="$DIRS $dir/fftw3"
> + done
> + for dir in $DIRS; do
> + # libraries might have different file extensions
> + for libext in a so dylib; do
> + # libraries can be in /lib or /lib64
> + for libdir in lib64 lib/x86_64-linux-gnu lib
> lib/i386-linux-gnu; do
> + FILES="include/fftw3.h $libdir/libfftw3.$libext"
> + # assume this is the one and check all needed files
> + FFTW3_DIR="$dir"
> + for file in $FILES; do
> + # discard this directory if one file was not found
> + if [ ! -r "$dir/$file" ]; then
> + unset FFTW3_DIR
> + break
> + fi
> + done
> + # don't look further if all files have been found
> + if [ -n "$FFTW3_DIR" ]; then
> + break
> + fi
> + done
> + # don't look further if all files have been found
> + if [ -n "$FFTW3_DIR" ]; then
> break
> fi
> done
> + # don't look further if all files have been found
> if [ -n "$FFTW3_DIR" ]; then
> break
> fi
> done
> -
> +
> if [ -z "$FFTW3_DIR" ]; then
> echo "BEGIN MESSAGE"
> echo "FFTW3 not found"
>
> _______________________________________________
> Commits mailing list
> Commits(a)cactuscode.org
> http://cactuscode.org/mailman/listinfo/commits
>
--
Erik Schnetter <schnetter(a)cct.lsu.edu>
http://www.perimeterinstitute.ca/personal/eschnetter/
Present were: Philipp, Frank, Roland, Peter, Steve, Eloisa
ET release:
* begin running test suites on clusters, assigned testers on wiki page
https://docs.einsteintoolkit.org/et-docs/Release_Details
* please run tests at least once by next week Monday
* name contest for release, same page
Simfactory:
* when not specifying BUILD one should make sure that the installed
version of external libraries is used (ie configure.sh finds the
installed libraries)
GRMHD paper:
* submitted to arXiv
* need to clean up material page
(http://einsteintoolkit.org/publications/2013_MHD/) today
* Philipp will push changes to trunk to make it possible to run
magnetized tov and collapse (again)
* add node count, memory consumption and run time (help wanted)
Yours,
Roland