Hi
Please consider joining the weekly Einstein Toolkit phone call at
9:30 am US central time on Mondays. For details on how to connect
and what agenda items are to be discussed, use the link below.
https://docs.einsteintoolkit.org/et-docs/Main_Page#Weekly_Users_Call
--Steve
Do you have a backtrace?
-erik
On Sunday, March 22, 2015, Einstein Toolkit <
trac-noreply(a)einsteintoolkit.org> wrote:
> #1326: running loopcontrol on strange number of threads fails
>
------------------------------------+---------------------------------------
> Reporter: rhaas | Owner:
> Type: defect | Status: new
> Priority: minor | Milestone:
> Component: EinsteinToolkit thorn | Version:
> Resolution: | Keywords: LoopControl
>
------------------------------------+---------------------------------------
>
> Comment (by rhaas):
>
> This still happens even with current (Sun Mar 22 18:46:21 CET 2015)
trunk,
> though failure looks a bit different now:
> {{{
> actus_sim:
>
/data/rhaas/postdoc/gr/ET_trunk/configs/sim/build/LoopControl/loopcontrol.cc:312:
> T {anonymous}::divexact(T, T) [with T = int]: Assertion `i % j == 0'
> failed
> }}}
>
> --
> Ticket URL: <https://trac.einsteintoolkit.org/ticket/1326#comment:1>
> Einstein Toolkit <http://einsteintoolkit.org>
> The Einstein Toolkit
> _______________________________________________
> Trac mailing list
> Trac(a)einsteintoolkit.org
> http://lists.einsteintoolkit.org/mailman/listinfo/trac
>
--
Erik Schnetter <schnetter(a)cct.lsu.edu>
http://www.perimeterinstitute.ca/personal/eschnetter/
Hi Erik, Roland, all,
After our discussion on last week's telecon, I followed Roland's
instructions on how to get the branch which has changes to how Carpet
handles prolongation with respect to OpenMP. I reran my simple scaling
test on Stampede Skylake nodes using this branch of Carpet
(rhaas/openmp-tasks) to test the scalability.
Attached is a plot showing the speeds for a variety of number of nodes
and how the 48 threads are distributed on the nodes between MPI
processes and OpenMP threads. I did this for three versions of the
ETK. 1. Fresh checkout of ET_2017_06. 2. The ET_2017_06 with Carpet
switched to the rhaas/openmp-tasks (labelled "Test On") 3. Again with
the checkout from #2, but without the parameters to enable the new
prolongation code (labelled "Test Off"). The run speeds used were
grabbed at iteration 256 from Carpet::physical_time_per_hour. No IO or
regridding.
For 4 and 8 nodes (ie 192 and 384 cores), there wasn't much difference
between the 3 trials. However, for 16 and 24 nodes (768 and 1152
cores), we see some improvement in run speed (10-15%) for many choices
of distribution of threads, again with a slight preference for 8
ranks/node.
I also ran the previous test (not using the openmp-tasks branch) on
comet, and found similar results as before.
Thanks,
Jim
On 01/21/2018 01:07 PM, Erik Schnetter wrote:
> James
>
> I looked at OpenMP performance in the Einstein Toolkit a few months
> ago, and I found that Carpet's prolongation operators are not well
> parallelized. There is a branch in Carpet (and a few related thorns)
> that apply a different OpenMP parallelization strategy, which seems to
> be more efficient. We are currently looking into cherry-picking the
> relevant changes from this branch (there are also many unrelated
> changes, since I experimented a lot) and putting them back into the
> master branch.
>
> These changes only help with prolongation, which seems to be a major
> contributor to non-OpenMP-scalability. I experimented with other
> changes as well. My findings (unfortunately without good solutions so
> far) are:
>
> - The standard OpenMP parallelization of loops over grid functions is
> not good for data cache locality. I experimented with padding arrays,
> ensuring that loop boundaries align with cache line boundaries, etc.,
> but this never worked quite satisfactorily -- MPI parallelization is
> still faster than OpenMP. In effect, the only reason one would use
> OpenMP is once one encounters MPI's scalability limits, so that
> OpenMP's non-scalability is less worse.
>
> - We could overlap calculations with communication. To do so, I have
> experimental changes that break loops over grid functions into tiles.
> Outer tiles need to wait for communication (synchronization or
> parallelization) to finish, while inner tiles can be calculated right
> away. Unfortunately, OpenMP does not support open-ended threads like
> this, so I'm using Qthreads <https://github.com/Qthreads/qthreads> and
> FunHPC <https://bitbucket.org/eschnett/funhpc.cxx> for this. The
> respective changes to Carpet, the scheduler, and thorns are
> significant, and I couldn't prove any performance improvements yet.
> However, once we removed other, more prominent non-scalability causes,
> I hope that this will become interesting.
>
> I haven't been attending the ET phone calls recently because Monday
> mornings aren't good for me schedule-wise. If you are interested, then
> we can ensure that we both attend at the same time and then discuss
> this. We need to make sure the Roland Haas is then also attending.
>
> -erik
>
>
> On Sat, Jan 20, 2018 at 10:21 AM, James Healy <jchsma(a)rit.edu
> <mailto:jchsma@rit.edu>> wrote:
>
> Hello all,
>
> I am trying to run on the new skylake processors on Stampede2 and
> while the run speeds we are obtaining are very good, we are
> concerned that we aren't optimizing properly when it comes to
> OpenMP. For instance, we see the best speeds when we use 8 MPI
> processors per node (with 6 threads each for a total of 48 total
> threads/node). Based on the architecture, we were expecting to
> see the best speeds with 2 MPI/node. Here is what I have tried:
>
> 1. Using the simfactory files for stampede2-skx (config file, run
> and submit scripts, and modules loaded) I compiled a version
> of ET_2017_06 using LazEv (RIT's evolution thorn) and
> McLachlan and submitted a series of runs that change both the
> number of nodes used, and how I distribute the 48 threads/node
> between MPI processes.
> 2. I use a standard low resolution grid, with no IO or
> regridding. Parameter file attached.
> 3. Run speeds are measured from Carpet::physical_time_per_hour at
> iteration 256.
> 4. I tried both with and without hwloc/SystemTopology.
> 5. For both McLachlan and LazEv, I see similar results, with 2
> MPI/node giving the worst results (see attached plot for
> McLachlan) and a slight preferences for 8 MPI/node.
>
> So my questions are:
>
> 1. Has there been any tests run by any other users on stampede2 skx?
> 2. Should we expect 2 MPI/node to be the optimal choice?
> 3. If so, are there any other configurations we can try that
> could help optimize?
>
> Thanks in advance!
>
> Jim Healy
>
>
> _______________________________________________
> Users mailing list
> Users(a)einsteintoolkit.org <mailto:Users@einsteintoolkit.org>
> http://lists.einsteintoolkit.org/mailman/listinfo/users
> <http://lists.einsteintoolkit.org/mailman/listinfo/users>
>
>
>
>
> --
> Erik Schnetter <schnetter(a)cct.lsu.edu <mailto:schnetter@cct.lsu.edu>>
> http://www.perimeterinstitute.ca/personal/eschnetter/
>
All,
Georgia Tech is hosting the Einstein Toolkit Workshop this year, and we are working on finalizing the dates. Below, I have linked to a when2meet poll for everyone to indicate their availability. All the dates are in June and we are aiming to have a 3 day workshop. We would like to finalize the dates before next Monday, so the final day to vote will be Sunday, February 4.
https://www.when2meet.com/?6644007-rVgmz
Thank you,
Deborah Ferguson
-----------------------------------------
Deborah Ferguson
Graduate Student
Georgia Institute of Technology
Center for Relativistic Astrophysics
Boggs 1-66
Hello all,
I am getting odd compile errors on OSX using osx-homebrew.cfg (in
master).
There seem to be two different errors right now:
* _mm_add_ps is not found (even though xmmintrin.h exists)
* ranlib (it uses XCode's) complains about the .a files that it claims
have no symbols. Forcing the use of ranlib from homebrew does not
show the error but then I need to hard-code the full path since
homebrew does not put ranlib in /usr/local/bin
I attach the make output (both types of errors are visible at the end).
A side node: to have remote commands work on OSX (ie remote compiling
*on* my OSX laptop not from it) only works if I add
"source /etc/profile" to my envsetup (otherwise /usr/local/bin is not
added to $PATH it seems).
Has anyone seen these types of errors before and / or could point me to
a "canonical" location for ranlib on OSX (rather
than /usr/local/Cellar/binutils/2.30/x86_64-apple-darwin17.3.0/bin/ranlib)?
Yours,
Roland
--
My email is as private as my paper mail. I therefore support encrypting
and signing email messages. Get my PGP key from http://pgp.mit.edu .
Present: Roland, Antoni Ramos Buades, Eloisa, Erik, Ian, Jim, Sam,
Steve, Peter, Yosef, Bhavesh, Bill, Carlos, Gabrielle, Muricia
Suarez, Roberto de Pietri, Qian, two anonymous callers
Invitation to join cosmo working groups
* if interested, please join the mailing list as outlined in the email
by Helvi at:
http://lists.einsteintoolkit.org/pipermail/users/2018-January/006021.html
Failing tests in PITTNullCode:
* Yosef tracked it down to a compiler bug. The only affected code is
the test, regular code in the thorn does not use the failing
combination of operations. A bug report has been filed with the gcc
developers
* would like to find out which versions are affected
* will follow Yosef's suggestion and remove the test
Release status:
* test are in the process of running on clusters
* have to find out (again) how to populate
http://einsteintoolkit.org/testsuite_results/index.php
* had release status meeting on Friday and could clarify most issues
* details are here:
https://docs.einsteintoolkit.org/et-docs/Release_Details
* Ian created tickets for the otustanding tasks (marked with the
milestone) in the trac system
* code name "Tesla"
OpenMP scaling:
* Jim ran openmp-tasks version on Stamped2 Skylake
* running on larger number of cores (~24 nodes) one sees an improved
runspeed using the task based prolongation compared to either the old
code or the new code with the options turned off
* this test used 11 reflevels and fairly high resolution
* Erik points out the importance of thread binding in Cactus which
requires the SystemTopology thorn. Jim will re-run with that thorn
active (~256 iterations is enough). Should also post the stdout files
where Carpet lists thread binding results
ET US meeting at GT:
* dates have still to be finalized
Yours,
Roland
--
My email is as private as my paper mail. I therefore support encrypting
and signing email messages. Get my PGP key from http://pgp.mit.edu .
Hello Chia-Hui,
this looks to me as if you are either not using the correct allocation
or are out of allocation time. Given that salloc worked for you I would
think is is the former.
Please check your simfactory/etc/defs.local.ini file in the [edison]
section to make sure that you have an
allocation = XXX
line where XXX is your allocation.
You can also check the SubmitScript (=SLURM) script used to see exactly
what simfactory passed to sbatch. The exact location is in the screen
output in the "Submit script:" line.
Yours,
Roland
> Dear Roland and whom it may concern,
> Before I set the submission as an interactive job which means I use the command :
> salloc -N 2 -p debug -L SCRATCH
> as I ran the example code.
> Recently I tried to set it as a batch job , so the command becomes:
> sbatch @SCRIPTFILE@
> However I met another error which seems the NIM information for me is not available as showed in the attached file.
> I am wondering whether it is related to the previous problem which we have discussed or there is something I was missing.
>
> Best regards,
> Chia-Hui Lin
> [cid:96df1161-586b-41e0-9ab7-a71b981bac7c]
> ________________________________
> 寄件者: ian.hinder(a)aei.mpg.de <ian.hinder(a)aei.mpg.de>
> 寄件日期: 2018年1月23日 下午 04:17:35
> 收件者: Roland Haas
> 副本: Einstein Toolkit Users; 林家暉
> 主旨: Re: [Users] questions about compilation
>
>
>
> On 21 Jan 2018, at 17:16, Roland Haas <rhaas(a)illinois.edu<mailto:rhaas@illinois.edu>> wrote:
>
> Hello Ian,
>
> Shouldn't these changes be backported to the current release?
> Sure. Would only be for a couple of weeks though. I am happy to
> backport them though.
>
> That would allow the machine to be used for science now by people just running "git pull" in the simfactory directory, rather than waiting until the end of February for the next release.
>
> --
> Ian Hinder
> http://members.aei.mpg.de/ianhin
>
--
My email is as private as my paper mail. I therefore support encrypting
and signing email messages. Get my PGP key from http://keys.gnupg.net.
We will have a release coordination meeting on Friday, 9 am Central
time, 10 am Easter.
when: Friday, January 26, 10:00 EST
where: https://bluejeans.com/510865816
We will discuss open tickets relative to the release. The relevant
queries can be found at https://trac.einsteintoolkit.org/
--Steve