Kentaro
When applying performance optimisations, one has to be careful not to be guided by one's experience, but rather only by measurements. Things behave often in a quite surprising manner. Do you have performance data to support your statements that array expressions are faster than do loops (without OpenMP)? Do you also have performance data to support your statement that OpenMP is slowing things down in this case? With performance data I refer e.g. to Cactus timer output for dissipation for a reasonable setup, running on a reliable machine (e.g. an HPC node).
My personal hypothesis would be that one of the changes you introduce (remove do loops, remove OpenMP parallelization) inhibits some compiler optimisation and thus leads to more consistent results.
As Frank says, reduction operations are not an issue here. I don't see which compiler optimisations would lead to random differences.
Do you have two versions of the produced executable, one that produces random changes and one that doesn't? Could you make these available to me? I would like to compare the produced machine code to see the difference.
-erik