-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
I'm doing a Cactus run which I've compiled with OpenMp support on my linux machine, and looking at the performance in the routine in which I explicitly added a "parallel do" directive. When I change the environment variable OMP_NUM_THREADS to anything larger than 1, it actually takes LONGER than the 1-thread run. I have *24* processors on my machine. It runs fine on my mac, but not my linux machine.
Does anyone have a suggestion?
I notice my version of gcc doesn't have --enable-openmp (see below)... but then the worst it should do is not have any effect when OMP_NUM_THREADS is set to anything.
Thanks! Scott
mpicc -v
Using built-in specs. Target: x86_64-linux-gnu Configured with: ../src/configure -v --with-pkgversion='Debian 4.4.5-8' --with-bugurl=file:///usr/share/doc/gcc-4.4/README.Bugs --enable-languages=c,c++,fortran,objc,obj-c++ --prefix=/usr --program-suffix=-4.4 --enable-shared --enable-multiarch --enable-linker-build-id --with-system-zlib --libexecdir=/usr/lib --without-included-gettext --enable-threads=posix --with-gxx-include-dir=/usr/include/c++/4.4 --libdir=/usr/lib --enable-nls --enable-clocale=gnu --enable-libstdcxx-debug --enable-objc-gc --with-arch-32=i586 --with-tune=generic --enable-checking=release --build=x86_64-linux-gnu --host=x86_64-linux-gnu --target=x86_64-linux-gnu Thread model: posix gcc version 4.4.5 (Debian 4.4.5-8)
libgomp is version 4.4.5-8 ... but there is no lib64gomp. Odd. I'm using Debian squeeze which supposedly contains lib64gomp.
- -- Scott H. Hawley, Ph.D. Asst. Prof. of Physics Chemistry & Physics Dept Office: Hitch 100D Belmont University Tel: +1-615-460-6206 Nashville, TN 37212 USA Fax: +1-615-460-5458 PGP Key at http://sks-keyservers.net
No doubt someone will ask for the code itself. The relevant part is given below. ( In old-school Fortran77) Specifically, the 'problem' I'm noticing is that the "% done" messages appear with less frequency and with lesser increment per wall clock time with OMP_NUM_THREADS > 1 than for OMP_NUM_THREADS = 1. The cpus get used alot more --- 'top' shows up to 2000% cpu usage for 24 threads --- but the wallclock time doesn't decrease at all.
Also note that whether I use the long OMP directive shown (with the 'shared' declarations and schedule, etc) and the 'END PARALLEL DO' at the end, or if I just use a simple '!$OMP PARALLEL DO' and *nothing else*, the execution time is *identical*.
Thanks again! -Scott
chunk = 8 do k = 1, nz write(msg,*)'[1F setbkgrnd: ' // & 'nz = ',nz,', % done = ',int(k*1.0d2/nz),' ' call writemessage(msg)
!$OMP PARALLEL DO SHARED(mask,agxx,agxy,agxz,agyy,agyz,agzz, !$OMP& aKxx,aKxy,aKxz,aKyy,aKyz,aKzz,ax,ay,az), !$OMP& SCHEDULE(STATIC,chunk) PRIVATE(j) do j = 1, ny do i = 1, nx
c if (ltrace) then c write(msg,*) '---------------' c call writemessage(msg) c endif
if (mask(i,j,k) .ne. m_ex .and. & (ibonly .eq. 0 .or. & mask(i,j,k) .eq. m_ib)) then
x = ax(i) y = ay(j) z = az(k)
c the following two include files just perform many pointwise calc's include 'gd.inc' include 'kd.inc'
agxx(i,j,k) = gxx agxy(i,j,k) = gxy agxz(i,j,k) = gxz agyy(i,j,k) = gyy agyz(i,j,k) = gyz agzz(i,j,k) = gzz
aKxx(i,j,k) = Kxx aKxy(i,j,k) = Kxy aKxz(i,j,k) = Kxz aKyy(i,j,k) = Kyy aKyz(i,j,k) = Kyz aKzz(i,j,k) = Kzz
else if (mask(i,j,k) .eq. m_ex) then c Excised points agxx(i,j,k) = exval agxy(i,j,k) = exval agxz(i,j,k) = exval agyy(i,j,k) = exval agyz(i,j,k) = exval agzz(i,j,k) = exval
aKxx(i,j,k) = exval aKxy(i,j,k) = exval aKxz(i,j,k) = exval aKyy(i,j,k) = exval aKyz(i,j,k) = exval aKzz(i,j,k) = exval endif
enddo enddo !$OMP END PARALLEL DO enddo
Scott
Thanks for showing the code.
I can think of several things that could go wrong:
1. It takes some time to start up and shut down parallel threads. Therefore, people usually parallelise the outermost loop, i.e. the k loop in your case. Parallelising the j loop requires starting and stopping threads nz times, which adds overhead.
2. I notice that you don't declare private variables. Variables are shared by default, and any local variables that you use inside the parallel region (and which are not arrays where you only access one element) need to be declared as private. In your case, these are probably x, y, z, gxx, gxy, gxz, etc. Did you compare results between serial and parallel runs? I would expect the results to differ, i.e. the current parallel code seems to have a serious error.
3. How much computational work is done inside this loop? If most of the time is spent in memory access writing to the gij and Kij arrays, then OpenMP won't be able to help. Only if there is sufficient computation going on will you see a benefit.
4. Since you say that you have 24 cores, I assume you have an AMD system. In this case, your machine consists of 4 subsystems that have 6 cores each, and communication between these 4 subsystems will be much slower than within each of these subsystems. People usually recommend to use not more than 6 OpenMP threads, and to ensure that these run within one of these subsystems. You can try setting the environment variable GOMP_CPU_AFFINITY='0-5' to force your threads to run on cores 0 to 5.
-erik
On Tue, May 17, 2011 at 1:00 PM, Scott Hawley scott.hawley@belmont.edu wrote:
No doubt someone will ask for the code itself. The relevant part is given below. ( In old-school Fortran77) Specifically, the 'problem' I'm noticing is that the "% done" messages appear with less frequency and with lesser increment per wall clock time with OMP_NUM_THREADS > 1 than for OMP_NUM_THREADS = 1. The cpus get used alot more --- 'top' shows up to 2000% cpu usage for 24 threads --- but the wallclock time doesn't decrease at all. Also note that whether I use the long OMP directive shown (with the 'shared' declarations and schedule, etc) and the 'END PARALLEL DO' at the end, or if I just use a simple '!$OMP PARALLEL DO' and *nothing else*, the execution time is *identical*. Thanks again! -Scott
chunk = 8 do k = 1, nz write(msg,*)'[1F setbkgrnd: ' // & 'nz = ',nz,', % done = ',int(k*1.0d2/nz),' ' call writemessage(msg)
!$OMP PARALLEL DO SHARED(mask,agxx,agxy,agxz,agyy,agyz,agzz, !$OMP& aKxx,aKxy,aKxz,aKyy,aKyz,aKzz,ax,ay,az), !$OMP& SCHEDULE(STATIC,chunk) PRIVATE(j) do j = 1, ny do i = 1, nx c if (ltrace) then c write(msg,*) '---------------' c call writemessage(msg) c endif
if (mask(i,j,k) .ne. m_ex .and. & (ibonly .eq. 0 .or. & mask(i,j,k) .eq. m_ib)) then x = ax(i) y = ay(j) z = az(k) c the following two include files just perform many pointwise calc's include 'gd.inc' include 'kd.inc' agxx(i,j,k) = gxx agxy(i,j,k) = gxy agxz(i,j,k) = gxz agyy(i,j,k) = gyy agyz(i,j,k) = gyz agzz(i,j,k) = gzz
aKxx(i,j,k) = Kxx aKxy(i,j,k) = Kxy aKxz(i,j,k) = Kxz aKyy(i,j,k) = Kyy aKyz(i,j,k) = Kyz aKzz(i,j,k) = Kzz
else if (mask(i,j,k) .eq. m_ex) then c Excised points agxx(i,j,k) = exval agxy(i,j,k) = exval agxz(i,j,k) = exval agyy(i,j,k) = exval agyz(i,j,k) = exval agzz(i,j,k) = exval
aKxx(i,j,k) = exval aKxy(i,j,k) = exval aKxz(i,j,k) = exval aKyy(i,j,k) = exval aKyz(i,j,k) = exval aKzz(i,j,k) = exval endif enddo enddo !$OMP END PARALLEL DO enddo
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Erik, Thanks for your ideas. #3 may be the most significant.
1. I switched the OMP directives to the outer loop, with the main result being, of course, that the "% done" line skips around, but NO change in execution speed.
2. I also increased the number of private variables as shown below. Again no change in speed. And by this I mean: 1 thread - the routine takes 11.3 seconds 2 threads - the routine takes 47.7 seconds 4 threads - the routine takes 40.1 seconds
These results are using code at the beginning of the loops which now reads... !$OMP PARALLEL DO SHARED(mask,agxx,agxy,agxz,agyy,agyz,agzz, !$OMP& aKxx,aKxy,aKxz,aKyy,aKyz,aKzz,ax,ay,az), !$OMP& SCHEDULE(STATIC,chunk) PRIVATE(k,j,i,gxx,gxy,gxz,gyy,gyz,gzz, !$OMP& x, y, z) do k = 1, nz write(msg,*)'[1F setbkgrnd: ' // & 'nz = ',nz,', % done = ',int(k*1.0d2/nz),' ' call writemessage(msg)
do j = 1, ny do i = 1, nx ...
3. There is ALOT of computational work 'per point' done, at each value of i,j,k; there are long formulas in the include files include 'gd.inc' include 'kd.inc' Wait...these include files contain *tons* of temporary variables that maybe should be private -- they were generated by maple, variables like 't1' thru 't250'. Do these all be need to be declared as private? i certainly don't want the various processors overwriting each others' work, which might be what they're doing. -- maybe they're even generating NaNs which would slow things down a bit!
4. Yea, even 2 threads vs one thread is a significant slowdown, as noted above.
5. Earlier I misspoke: "It works fine on my mac" means OpenMP works fine on my Mac, for a *different* program I wrote. This program *also* works fine on my linux box. So it's *very likely* that my current issue is "user error", and not a misconfigured OpenMP. lol.
Thanks, Scott
-- Scott H. Hawley, Ph.D. Asst. Prof. of Physics Chemistry & Physics Dept Office: Hitch 100D Belmont University Tel: +1-615-460-6206 Nashville, TN 37212 USA Fax: +1-615-460-5458 PGP Key at http://sks-keyservers.net
On May 17, 2011, at 12:46 PM, Erik Schnetter wrote:
Scott
Thanks for showing the code.
I can think of several things that could go wrong:
- It takes some time to start up and shut down parallel threads.
Therefore, people usually parallelise the outermost loop, i.e. the k loop in your case. Parallelising the j loop requires starting and stopping threads nz times, which adds overhead.
- I notice that you don't declare private variables. Variables are
shared by default, and any local variables that you use inside the parallel region (and which are not arrays where you only access one element) need to be declared as private. In your case, these are probably x, y, z, gxx, gxy, gxz, etc. Did you compare results between serial and parallel runs? I would expect the results to differ, i.e. the current parallel code seems to have a serious error.
- How much computational work is done inside this loop? If most of
the time is spent in memory access writing to the gij and Kij arrays, then OpenMP won't be able to help. Only if there is sufficient computation going on will you see a benefit.
- Since you say that you have 24 cores, I assume you have an AMD
system. In this case, your machine consists of 4 subsystems that have 6 cores each, and communication between these 4 subsystems will be much slower than within each of these subsystems. People usually recommend to use not more than 6 OpenMP threads, and to ensure that these run within one of these subsystems. You can try setting the environment variable GOMP_CPU_AFFINITY='0-5' to force your threads to run on cores 0 to 5.
-erik
On Tue, May 17, 2011 at 1:00 PM, Scott Hawley scott.hawley@belmont.edu wrote:
No doubt someone will ask for the code itself. The relevant part is given below. ( In old-school Fortran77) Specifically, the 'problem' I'm noticing is that the "% done" messages appear with less frequency and with lesser increment per wall clock time with OMP_NUM_THREADS > 1 than for OMP_NUM_THREADS = 1. The cpus get used alot more --- 'top' shows up to 2000% cpu usage for 24 threads --- but the wallclock time doesn't decrease at all. Also note that whether I use the long OMP directive shown (with the 'shared' declarations and schedule, etc) and the 'END PARALLEL DO' at the end, or if I just use a simple '!$OMP PARALLEL DO' and *nothing else*, the execution time is *identical*. Thanks again! -Scott
chunk = 8 do k = 1, nz write(msg,*)'[1F setbkgrnd: ' // & 'nz = ',nz,', % done = ',int(k*1.0d2/nz),' ' call writemessage(msg)!$OMP PARALLEL DO SHARED(mask,agxx,agxy,agxz,agyy,agyz,agzz, !$OMP& aKxx,aKxy,aKxz,aKyy,aKyz,aKzz,ax,ay,az), !$OMP& SCHEDULE(STATIC,chunk) PRIVATE(j) do j = 1, ny do i = 1, nx c if (ltrace) then c write(msg,*) '---------------' c call writemessage(msg) c endif
if (mask(i,j,k) .ne. m_ex .and. & (ibonly .eq. 0 .or. & mask(i,j,k) .eq. m_ib)) then x = ax(i) y = ay(j) z = az(k)c the following two include files just perform many pointwise calc's include 'gd.inc' include 'kd.inc' agxx(i,j,k) = gxx agxy(i,j,k) = gxy agxz(i,j,k) = gxz agyy(i,j,k) = gyy agyz(i,j,k) = gyz agzz(i,j,k) = gzz
aKxx(i,j,k) = Kxx aKxy(i,j,k) = Kxy aKxz(i,j,k) = Kxz aKyy(i,j,k) = Kyy aKyz(i,j,k) = Kyz aKzz(i,j,k) = Kzz else if (mask(i,j,k) .eq. m_ex) thenc Excised points agxx(i,j,k) = exval agxy(i,j,k) = exval agxz(i,j,k) = exval agyy(i,j,k) = exval agyz(i,j,k) = exval agzz(i,j,k) = exval
aKxx(i,j,k) = exval aKxy(i,j,k) = exval aKxz(i,j,k) = exval aKyy(i,j,k) = exval aKyz(i,j,k) = exval aKzz(i,j,k) = exval endif enddo enddo!$OMP END PARALLEL DO enddo
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Erik Schnetter schnetter@cct.lsu.edu http://www.cct.lsu.edu/~eschnett/
Hi,
On Tue, May 17, 2011 at 03:02:00PM -0700, Scott Hawley wrote:
Do these all be need to be declared as private?
If the temporary variables are declared only inside the loop they are automatically thread-local. Oh wait, that is Fortran. Well - in that case you should either declare all private, or (maybe easier) put the include files into separate functions, declare the temporary variables only there and call the functions from within the loop, in which case they also don't have to be specified for openmp (as long as they are not static).
i certainly don't want the various processors overwriting each others' work, which might be what they're doing. -- maybe they're even generating NaNs which would slow things down a bit!
You should see that in the results though. It might make sense to first make sure that the results with different numbers of threads are the same (depending on the problem you might actually get bit-by-bit identical results), and work on optimization later. I agree that your slow-down actually points towards some bug.
Frank
Hi Scott,
On Tue, 17 May 2011, Frank Loeffler wrote:
Hi,
On Tue, May 17, 2011 at 03:02:00PM -0700, Scott Hawley wrote:
Do these all be need to be declared as private?
If the temporary variables are declared only inside the loop they are automatically thread-local. Oh wait, that is Fortran. Well - in that case you should either declare all private, or (maybe easier) put the include files into separate functions, declare the temporary variables only there and call the functions from within the loop, in which case they also don't have to be specified for openmp (as long as they are not static).
i certainly don't want the various processors overwriting each others' work, which might be what they're doing. -- maybe they're even generating NaNs which would slow things down a bit!
Alternatively you may use the DEFAULT(PRIVATE) clause, so that you only have to specify the shared variables. However, in that case you have to make sure to really declare all the shared variables as shared, since otherwise all processors will have to allocate storage and if they are 3d variables this will slow down the code and increase memory consumption. Also private variables have undefined values on entry to the parallel region so not declaring all shared variables properly can also adversely affect the result. So be careful.
You should see that in the results though. It might make sense to first make sure that the results with different numbers of threads are the same (depending on the problem you might actually get bit-by-bit identical results), and work on optimization later. I agree that your slow-down actually points towards some bug.
Frank
Cheers,
Peter
Erik, Frank, Peter: Thanks guys. I will pursue your suggestions.
-- Scott H. Hawley, Ph.D. Asst. Prof. of Physics Chemistry & Physics Dept Office: Hitch 100D Belmont University Tel: +1-615-460-6206 Nashville, TN 37212 USA Fax: +1-615-460-5458 PGP Key at http://sks-keyservers.net
On May 18, 2011, at 12:14 AM, Peter Diener wrote:
Hi Scott,
On Tue, 17 May 2011, Frank Loeffler wrote:
Hi,
On Tue, May 17, 2011 at 03:02:00PM -0700, Scott Hawley wrote:
Do these all be need to be declared as private?
If the temporary variables are declared only inside the loop they are automatically thread-local. Oh wait, that is Fortran. Well - in that case you should either declare all private, or (maybe easier) put the include files into separate functions, declare the temporary variables only there and call the functions from within the loop, in which case they also don't have to be specified for openmp (as long as they are not static).
i certainly don't want the various processors overwriting each others' work, which might be what they're doing. -- maybe they're even generating NaNs which would slow things down a bit!
Alternatively you may use the DEFAULT(PRIVATE) clause, so that you only have to specify the shared variables. However, in that case you have to make sure to really declare all the shared variables as shared, since otherwise all processors will have to allocate storage and if they are 3d variables this will slow down the code and increase memory consumption. Also private variables have undefined values on entry to the parallel region so not declaring all shared variables properly can also adversely affect the result. So be careful.
You should see that in the results though. It might make sense to first make sure that the results with different numbers of threads are the same (depending on the problem you might actually get bit-by-bit identical results), and work on optimization later. I agree that your slow-down actually points towards some bug.
Frank
Cheers,
Peter
Ok. Still having problems. I defaulted everything to private and explicitly declared my shared's. Now what happens is that the outer "k" loop never gets incremented. Even if I run with only one thread, "k" always equals 1.
So the snippet of code follows. Where m_ex and m_ib are parameters in Fortran and are hard-coded as numbers by the compiler. If I compile with -fopenmp it works fine, but at the -fopenmp and "setenv OMP_NUM_THREADS 1" and it won't increment.
Any new ideas? Thanks in advance. -Scott
!$OMP PARALLEL DO DEFAULT(PRIVATE) SHARED(mask,ibonly,ax,ay,az, !$OMP& agxx,agxy,agxz,agyy,agyz,agzz, !$OMP& aKxx,aKxy,aKxz,aKyy,aKyz,aKzz,nx,ny,nz) !$OMP& SCHEDULE(STATIC,chunk) do k = 1, nz write(msg,*)' setbkgrnd: ' // & 'nz = ',nz,', % done = ',int(k*1.0d2/nz),' ' call writemessage(msg)
do j = 1, ny do i = 1, nx
if (mask(i,j,k) .ne. m_ex .and. & (ibonly .eq. 0 .or. & mask(i,j,k) .eq. m_ib)) then
x = ax(i) y = ay(j) z = az(k) include 'gd.inc' include 'kd.inc'
agxx(i,j,k) = gxx agxy(i,j,k) = gxy agxz(i,j,k) = gxz agyy(i,j,k) = gyy agyz(i,j,k) = gyz agzz(i,j,k) = gzz
aKxx(i,j,k) = Kxx aKxy(i,j,k) = Kxy aKxz(i,j,k) = Kxz aKyy(i,j,k) = Kyy aKyz(i,j,k) = Kyz aKzz(i,j,k) = Kzz else if (mask(i,j,k) .eq. m_ex) then c Excised points agxx(i,j,k) = exval agxy(i,j,k) = exval agxz(i,j,k) = exval agyy(i,j,k) = exval agyz(i,j,k) = exval agzz(i,j,k) = exval
aKxx(i,j,k) = exval aKxy(i,j,k) = exval aKxz(i,j,k) = exval aKyy(i,j,k) = exval aKyz(i,j,k) = exval aKzz(i,j,k) = exval endif
enddo enddo enddo
There is no explicit OMP END DO statement because it's optional.
-- Scott H. Hawley, Ph.D. Asst. Prof. of Physics Chemistry & Physics Dept Office: Hitch 100D Belmont University Tel: +1-615-460-6206 Nashville, TN 37212 USA Fax: +1-615-460-5458 PGP Key at http://sks-keyservers.net
On May 18, 2011, at 4:37 PM, Scott Hawley wrote:
Erik, Frank, Peter: Thanks guys. I will pursue your suggestions.
-- Scott H. Hawley, Ph.D. Asst. Prof. of Physics Chemistry & Physics Dept Office: Hitch 100D Belmont University Tel: +1-615-460-6206 Nashville, TN 37212 USA Fax: +1-615-460-5458 PGP Key at http://sks-keyservers.net
On May 18, 2011, at 12:14 AM, Peter Diener wrote:
Hi Scott,
On Tue, 17 May 2011, Frank Loeffler wrote:
Hi,
On Tue, May 17, 2011 at 03:02:00PM -0700, Scott Hawley wrote:
Do these all be need to be declared as private?
If the temporary variables are declared only inside the loop they are automatically thread-local. Oh wait, that is Fortran. Well - in that case you should either declare all private, or (maybe easier) put the include files into separate functions, declare the temporary variables only there and call the functions from within the loop, in which case they also don't have to be specified for openmp (as long as they are not static).
i certainly don't want the various processors overwriting each others' work, which might be what they're doing. -- maybe they're even generating NaNs which would slow things down a bit!
Alternatively you may use the DEFAULT(PRIVATE) clause, so that you only have to specify the shared variables. However, in that case you have to make sure to really declare all the shared variables as shared, since otherwise all processors will have to allocate storage and if they are 3d variables this will slow down the code and increase memory consumption. Also private variables have undefined values on entry to the parallel region so not declaring all shared variables properly can also adversely affect the result. So be careful.
You should see that in the results though. It might make sense to first make sure that the results with different numbers of threads are the same (depending on the problem you might actually get bit-by-bit identical results), and work on optimization later. I agree that your slow-down actually points towards some bug.
Frank
Cheers,
Peter
<PGP.sig><ATT00001..txt>
Hey Scott,
OpenMP has a huge overhead. It must create new functions, which then are started as threads, and this takes some time. During a OpenMP tutorial I have made some tests with a matrix vector multiplication. It turned out, that you can see an increase of the scalability using OpenMP in such a case only from a vectory size with more then 20000 elements.
If you are using OpenMP additional to MPI, and further if you use simfactory to start your runs, I have another answer.
Recently I have made some comparisons between activated OpenMP extensions and only MPI, and I have used simfactory to start my simulations. Simfactory has not done that, what I have expected. Let me explain what I mean.
If I activate OpenMP by setting OMP_NUM_THREADS to 4, and then use simfactory with --procs=16, then what simfactory finally makes a run with
-pe openmpi 16
but(!!)
numprocs=4 numthreads=4
Using only MPI you will have in such a case
numprocs=16 numthreads=1
Hope that helps. So have your program running on 16 nodes with MPI only, compared to 4 nodes and additional 4 OpenMP threads per node. In such a case MPI is always faster, if your code in scaling.
Hope that helps.
Cheers
Alexander
On Thursday, May 19, 2011 00:33:49 Scott Hawley wrote:
Ok. Still having problems. I defaulted everything to private and explicitly declared my shared's. Now what happens is that the outer "k" loop never gets incremented. Even if I run with only one thread, "k" always equals 1.
So the snippet of code follows. Where m_ex and m_ib are parameters in Fortran and are hard-coded as numbers by the compiler. If I compile with -fopenmp it works fine, but at the -fopenmp and "setenv OMP_NUM_THREADS 1" and it won't increment.
Any new ideas? Thanks in advance. -Scott
!$OMP PARALLEL DO DEFAULT(PRIVATE) SHARED(mask,ibonly,ax,ay,az, !$OMP& agxx,agxy,agxz,agyy,agyz,agzz, !$OMP& aKxx,aKxy,aKxz,aKyy,aKyz,aKzz,nx,ny,nz) !$OMP& SCHEDULE(STATIC,chunk) do k = 1, nz write(msg,*)' setbkgrnd: ' // & 'nz = ',nz,', % done = ',int(k*1.0d2/nz),' ' call writemessage(msg)
do j = 1, ny do i = 1, nx if (mask(i,j,k) .ne. m_ex .and. & (ibonly .eq. 0 .or. & mask(i,j,k) .eq. m_ib)) then x = ax(i) y = ay(j) z = az(k) include 'gd.inc' include 'kd.inc' agxx(i,j,k) = gxx agxy(i,j,k) = gxy agxz(i,j,k) = gxz agyy(i,j,k) = gyy agyz(i,j,k) = gyz agzz(i,j,k) = gzz aKxx(i,j,k) = Kxx aKxy(i,j,k) = Kxy aKxz(i,j,k) = Kxz aKyy(i,j,k) = Kyy aKyz(i,j,k) = Kyz aKzz(i,j,k) = Kzz else if (mask(i,j,k) .eq. m_ex) thenc Excised points agxx(i,j,k) = exval agxy(i,j,k) = exval agxz(i,j,k) = exval agyy(i,j,k) = exval agyz(i,j,k) = exval agzz(i,j,k) = exval
aKxx(i,j,k) = exval aKxy(i,j,k) = exval aKxz(i,j,k) = exval aKyy(i,j,k) = exval aKyz(i,j,k) = exval aKzz(i,j,k) = exval endif enddo enddo enddoThere is no explicit OMP END DO statement because it's optional.
-- Scott H. Hawley, Ph.D. Asst. Prof. of Physics Chemistry & Physics Dept Office: Hitch 100D Belmont University Tel: +1-615-460-6206 Nashville, TN 37212 USA Fax: +1-615-460-5458 PGP Key at http://sks-keyservers.net
On May 18, 2011, at 4:37 PM, Scott Hawley wrote:
Erik, Frank, Peter: Thanks guys. I will pursue your suggestions.
-- Scott H. Hawley, Ph.D. Asst. Prof. of Physics Chemistry & Physics Dept Office: Hitch 100D Belmont University Tel: +1-615-460-6206 Nashville, TN 37212 USA Fax: +1-615-460-5458 PGP Key at http://sks-keyservers.net
On May 18, 2011, at 12:14 AM, Peter Diener wrote:
Hi Scott,
On Tue, 17 May 2011, Frank Loeffler wrote:
Hi,
On Tue, May 17, 2011 at 03:02:00PM -0700, Scott Hawley wrote:
Do these all be need to be declared as private?
If the temporary variables are declared only inside the loop they are automatically thread-local. Oh wait, that is Fortran. Well - in that case you should either declare all private, or (maybe easier) put the include files into separate functions, declare the temporary variables only there and call the functions from within the loop, in which case they also don't have to be specified for openmp (as long as they are not static).
i certainly don't want the various processors overwriting each others' work, which might be what they're doing. -- maybe they're even generating NaNs which would slow things down a bit!
Alternatively you may use the DEFAULT(PRIVATE) clause, so that you only have to specify the shared variables. However, in that case you have to make sure to really declare all the shared variables as shared, since otherwise all processors will have to allocate storage and if they are 3d variables this will slow down the code and increase memory consumption. Also private variables have undefined values on entry to the parallel region so not declaring all shared variables properly can also adversely affect the result. So be careful.
You should see that in the results though. It might make sense to first make sure that the results with different numbers of threads are the same (depending on the problem you might actually get bit-by-bit identical results), and work on optimization later. I agree that your slow-down actually points towards some bug.
Frank
Cheers,
Peter
<PGP.sig><ATT00001..txt>
Thanks Alexander. The overhead associated with re-allocating all those temporary variables is the reason why I used include files instead of defining separate subroutines. My code works with MPI, but I was looking for a way to eek out a little more parallelism. The domain decomposition routine for MPI needs to be changed so I can get more than 8 nodes, and I thought OpenMP would be the way to do it.
And I believe I still can, *elsewhere* in the code. The particular routine I chose was just perhaps not a good one. I'm just removing the OpenMP directives from this routine, and let it only be parallel via MPI.
Thanks!
-Scott
-- Scott H. Hawley, Ph.D. Asst. Prof. of Physics Chemistry & Physics Dept Office: Hitch 100D Belmont University Tel: +1-615-460-6206 Nashville, TN 37212 USA Fax: +1-615-460-5458 PGP Key at http://sks-keyservers.net
On May 19, 2011, at 2:14 AM, Alexander Beck-Ratzka wrote:
Hey Scott,
OpenMP has a huge overhead. It must create new functions, which then are started as threads, and this takes some time. During a OpenMP tutorial I have made some tests with a matrix vector multiplication. It turned out, that you can see an increase of the scalability using OpenMP in such a case only from a vectory size with more then 20000 elements.
If you are using OpenMP additional to MPI, and further if you use simfactory to start your runs, I have another answer.
Recently I have made some comparisons between activated OpenMP extensions and only MPI, and I have used simfactory to start my simulations. Simfactory has not done that, what I have expected. Let me explain what I mean.
If I activate OpenMP by setting OMP_NUM_THREADS to 4, and then use simfactory with --procs=16, then what simfactory finally makes a run with
-pe openmpi 16
but(!!)
numprocs=4 numthreads=4
Using only MPI you will have in such a case
numprocs=16 numthreads=1
Hope that helps. So have your program running on 16 nodes with MPI only, compared to 4 nodes and additional 4 OpenMP threads per node. In such a case MPI is always faster, if your code in scaling.
Hope that helps.
Cheers
Alexander
On Thursday, May 19, 2011 00:33:49 Scott Hawley wrote:
Ok. Still having problems. I defaulted everything to private and explicitly declared my shared's. Now what happens is that the outer "k" loop never gets incremented. Even if I run with only one thread, "k" always equals 1.
So the snippet of code follows. Where m_ex and m_ib are parameters in Fortran and are hard-coded as numbers by the compiler. If I compile with -fopenmp it works fine, but at the -fopenmp and "setenv OMP_NUM_THREADS 1" and it won't increment.
Any new ideas? Thanks in advance. -Scott
!$OMP PARALLEL DO DEFAULT(PRIVATE) SHARED(mask,ibonly,ax,ay,az, !$OMP& agxx,agxy,agxz,agyy,agyz,agzz, !$OMP& aKxx,aKxy,aKxz,aKyy,aKyz,aKzz,nx,ny,nz) !$OMP& SCHEDULE(STATIC,chunk) do k = 1, nz write(msg,*)' setbkgrnd: ' // & 'nz = ',nz,', % done = ',int(k*1.0d2/nz),' ' call writemessage(msg)
do j = 1, ny do i = 1, nx if (mask(i,j,k) .ne. m_ex .and. & (ibonly .eq. 0 .or. & mask(i,j,k) .eq. m_ib)) then x = ax(i) y = ay(j) z = az(k) include 'gd.inc' include 'kd.inc' agxx(i,j,k) = gxx agxy(i,j,k) = gxy agxz(i,j,k) = gxz agyy(i,j,k) = gyy agyz(i,j,k) = gyz agzz(i,j,k) = gzz aKxx(i,j,k) = Kxx aKxy(i,j,k) = Kxy aKxz(i,j,k) = Kxz aKyy(i,j,k) = Kyy aKyz(i,j,k) = Kyz aKzz(i,j,k) = Kzz else if (mask(i,j,k) .eq. m_ex) thenc Excised points agxx(i,j,k) = exval agxy(i,j,k) = exval agxz(i,j,k) = exval agyy(i,j,k) = exval agyz(i,j,k) = exval agzz(i,j,k) = exval
aKxx(i,j,k) = exval aKxy(i,j,k) = exval aKxz(i,j,k) = exval aKyy(i,j,k) = exval aKyz(i,j,k) = exval aKzz(i,j,k) = exval endif enddo enddo enddoThere is no explicit OMP END DO statement because it's optional.
-- Scott H. Hawley, Ph.D. Asst. Prof. of Physics Chemistry & Physics Dept Office: Hitch 100D Belmont University Tel: +1-615-460-6206 Nashville, TN 37212 USA Fax: +1-615-460-5458 PGP Key at http://sks-keyservers.net
On May 18, 2011, at 4:37 PM, Scott Hawley wrote:
Erik, Frank, Peter: Thanks guys. I will pursue your suggestions.
-- Scott H. Hawley, Ph.D. Asst. Prof. of Physics Chemistry & Physics Dept Office: Hitch 100D Belmont University Tel: +1-615-460-6206 Nashville, TN 37212 USA Fax: +1-615-460-5458 PGP Key at http://sks-keyservers.net
On May 18, 2011, at 12:14 AM, Peter Diener wrote:
Hi Scott,
On Tue, 17 May 2011, Frank Loeffler wrote:
Hi,
On Tue, May 17, 2011 at 03:02:00PM -0700, Scott Hawley wrote:
Do these all be need to be declared as private?
If the temporary variables are declared only inside the loop they are automatically thread-local. Oh wait, that is Fortran. Well - in that case you should either declare all private, or (maybe easier) put the include files into separate functions, declare the temporary variables only there and call the functions from within the loop, in which case they also don't have to be specified for openmp (as long as they are not static).
i certainly don't want the various processors overwriting each others' work, which might be what they're doing. -- maybe they're even generating NaNs which would slow things down a bit!
Alternatively you may use the DEFAULT(PRIVATE) clause, so that you only have to specify the shared variables. However, in that case you have to make sure to really declare all the shared variables as shared, since otherwise all processors will have to allocate storage and if they are 3d variables this will slow down the code and increase memory consumption. Also private variables have undefined values on entry to the parallel region so not declaring all shared variables properly can also adversely affect the result. So be careful.
You should see that in the results though. It might make sense to first make sure that the results with different numbers of threads are the same (depending on the problem you might actually get bit-by-bit identical results), and work on optimization later. I agree that your slow-down actually points towards some bug.
Frank
Cheers,
Peter
<PGP.sig><ATT00001..txt>
-- +++++++++++++++++++++++++++++++++++++++++++++++++ Dr. Alexander Beck-Ratzka - team leader eScience group
MPI for Gravitational Physics (Albert Einstein Institute) Am Mühlenberg 1 D-14476 Potsdam
Tel.: 0049 -(0)331 - 567-7192 Email: alexander.beck-ratzka@aei.mpg.de +++++++++++++++++++++++++++++++++++++++++++++++++ _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On Wed, May 18, 2011 at 6:33 PM, Scott Hawley scott.hawley@belmont.edu wrote:
Ok. Still having problems. I defaulted everything to private and explicitly declared my shared's. Now what happens is that the outer "k" loop never gets incremented. Even if I run with only one thread, "k" always equals 1.
I believe that the I/O statements (the write statement) cannot be called from a parallel region. You need to enclose it in a critical statement or similar if you want to keep it. I would suggest instead to comment it out.
-erik
As a follow-up....
In terms of performance: it turns out that, despite the overhead associated with defining 1000's of temporary variables over & over, it's not actually that time-consuming. Thus, parallelization via OpenMP "wins" in the end.
In terms of coding: Trying to declare "privates" and/or "shared" carefully was too much work. In the end I put it all in another subroutine --- Frank's original suggestion -- and got the benefit of the compiler helping me track the variables. :-)
So the relecant part of my code now reads
!$OMP PARALLEL DO DEFAULT(SHARED) SCHEDULE(STATIC,chunk) PRIVATE(k) do k = 1, nz call bksslice( & agxx, agxy, agxz, agyy, agyz, agzz, & aKxx, aKxy, aKxz, aKyy, aKyz, aKzz, & mask, & M1,a1,pos1x,pos1y,pos1z,vx_1,vy_1,vz_1,v_1,v1,n1, & gam_1,gam1,theta1,phi1, & M2,a2,pos2x,pos2y,pos2z,vx_2,vy_2,vz_2,v_2,v2,n2, & gam_2,gam2,theta2,phi2, & ax, ay, az, nx, ny, nz, ibonly, k & ) enddo
And that's it. Also, 'bkslice' includes a write statement for the "progress indicator" I wanted to preserve. All works well.
FYI. Thanks guys. -Scott
users@lists.einsteintoolkit.org