Hi, all,
To Frank,
However using "-fp-model precise" in option list( i.e., this option is used by all fortran77 code), the run will become much slow down.
Out of interest: how much is "much"?
I didn't test it. However the document( p11 in http://software.intel.com/sites/default/files/article/164389/fp-consistency-... ) say that the performance is reduced by 12-15%.
To Peter,
Just to clarify regarding the OpenMP parallelization. Even though it's supposed to be true that Fortran90 array constructs can be OpenMP parallelized using the !$OMP WORKSHARE directive, it turns out that some compilers (most notably the intel compilers) doesn't actually parallelize these constructs. Instead one processor does all the work. For that unfortunate reason the loop construct is much more efficient than the array construct on many machines.
Thank you for telling me about it. I have never used "!$OMP WORKSHARE" directive. So I tested this directive in the following performance test.
To Erik,
Do you have performance data to support your statements that array expressions are faster than do loops (without OpenMP)?
Intel compiler document such as http://www-h.eng.cam.ac.uk/help/languages/fortran/intelfortran/for_ug2.pdf , say "Rather than use explicit loops for array access, use elemental array operations...(P20)". So I thought that array expressions are faster than do-loops, although I had never compared these.
I'm attaching data of performance test for attached par file. Due to these data, both expressions are almost same.
Do you also have performance data to support your statement that OpenMP is slowing things down in this case?
Please see the attached file. Apparently the case (4) is the fastest, that is, OMP palatalization shouldn't be used for these small section.
My personal hypothesis would be that one of the changes you introduce (remove do loops, remove OpenMP parallelization) inhibits some compiler optimisation and thus leads to more consistent results.
Already I mentioned, "introducing a local array var, which is a copy of the input variable" prevent from random results.
Do you have two versions of the produced executable, one that produces random changes and one that doesn't? Could you make these available to me? I would like to compare the produced machine code to see the difference.
I guess you can access damiana@AEI. You can use the following executable files in "/lustre/datura/takami/Share_Dir/EXE__Test_Dissipation/".
For random results, "cactus_GRHydro.F77--datura.OMP-NO_PRECISE-NO".
For consistent results, "cactus_GRHydro.F77--datura.OMP-NO_PRECISE-YES".
Here, both are exactly same source programs and compile options except for "-fp-model precise" option.
In order to compare the results, usually I use the following command: $h5diff \ OUTPUT_DATA_1/CHECKPOINT/checkpoint.chkpt.it_5.file_0.h5 \ OUTPUT_DATA_2/CHECKPOINT/checkpoint.chkpt.it_5.file_0.h5
Kentaro TAKAMI