Dear all,
I encountered the problem in which the simulation result is randomly different at each run( of course I used same executable, par file, HPC and so on.).
Finally I'm doubting the origin of the problem is "CactusNumerical/Dissipation", because when I don't activate this thorn, the problem can be disappeared.
In this thorn, "dissipation.c" calls the fortran subroutine "apply_dissipation.F77" directly without scheduling in CACTUS. I guess the variables between C and fortran become inconsistency in this stage.
What do you think? I don't have much knowledge to call fortran subroutine from directly c language. So I need comments from expert.
Kentaro TAKAMI
Hi Kentaro,
On Fri, Feb 22, 2013 at 04:48:48PM +0100, Kentaro Takami wrote:
I encountered the problem in which the simulation result is randomly different at each run( of course I used same executable, par file, HPC and so on.).
I see the same for some of my runs, but I wouldn't expect otherwise.
In my case the simulation depends on 'global' quantities, e.g., quantities obtained by reductions. Reductions cannot be done exactly, at least not efficiently. Thus, values obtained via reductions (can) have always an error that is different even when running the same simulation twice, at least when done in parallel. If your simulation depends on this, then these (small) differences can quickly grow, especially in iterative schemes, e.g., hydro.
Finally I'm doubting the origin of the problem is "CactusNumerical/Dissipation", because when I don't activate this thorn, the problem can be disappeared.
This is interesting. So, you say you don't see random changes when Dissipation isn't active, but you do when you activate it, and everything else is the same? Which variables do you apply dissipation to?
In this thorn, "dissipation.c" calls the fortran subroutine "apply_dissipation.F77" directly without scheduling in CACTUS. I guess the variables between C and fortran become inconsistency in this stage.
I don't see a problem in the code with this. Calling fortran from C the way done in this thorn seems to be ok.
Frank
On 22 Feb 2013, at 17:43, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi Kentaro,
On Fri, Feb 22, 2013 at 04:48:48PM +0100, Kentaro Takami wrote:
I encountered the problem in which the simulation result is randomly different at each run( of course I used same executable, par file, HPC and so on.).
I see the same for some of my runs, but I wouldn't expect otherwise.
In my case the simulation depends on 'global' quantities, e.g., quantities obtained by reductions. Reductions cannot be done exactly, at least not efficiently. Thus, values obtained via reductions (can) have always an error that is different even when running the same simulation twice, at least when done in parallel. If your simulation depends on this, then these (small) differences can quickly grow, especially in iterative schemes, e.g., hydro.
I disagree. Reductions should be deterministic; assuming the same number of MPI processes, the contributions to the reduction should always be added in the same order. If you change the number of MPI processes, then I agree that the result of reductions can change, as the order of the sum over points will change, and floating point addition is not associative.
The only "excuse" I have seen for the same simulation (exe, parfile, machine, etc) giving different results on different runs is related to a comment in http://software.intel.com/sites/default/files/article/164389/fp-consistency-... (section "Second Example from WRF ") which says that depending on the alignment of the address in memory of a particular loop, either a vectorised or a nonvectorised version of the loop may be called.
Kentaro:
What quantity is different, and by how much? What happens if you compile without optimisation? Are the differences present in the initial data, or just after some evolution steps? How many steps? Where in the grid do the differences appear?
Finally I'm doubting the origin of the problem is "CactusNumerical/Dissipation", because when I don't activate this thorn, the problem can be disappeared.
This is interesting. So, you say you don't see random changes when Dissipation isn't active, but you do when you activate it, and everything else is the same? Which variables do you apply dissipation to?
In this thorn, "dissipation.c" calls the fortran subroutine "apply_dissipation.F77" directly without scheduling in CACTUS. I guess the variables between C and fortran become inconsistency in this stage.
I don't see a problem in the code with this. Calling fortran from C the way done in this thorn seems to be ok.
On Fri, Feb 22, 2013 at 12:01 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 22 Feb 2013, at 17:43, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi Kentaro,
On Fri, Feb 22, 2013 at 04:48:48PM +0100, Kentaro Takami wrote:
I encountered the problem in which the simulation result is randomly
different
at each run( of course I used same executable, par file, HPC and so
on.).
I see the same for some of my runs, but I wouldn't expect otherwise.
In my case the simulation depends on 'global' quantities, e.g., quantities obtained by reductions. Reductions cannot be done exactly, at least not efficiently. Thus, values obtained via reductions (can) have always an error that is different even when running the same simulation twice, at least when done in parallel. If your simulation depends on this, then these (small) differences can quickly grow, especially in iterative schemes, e.g., hydro.
I disagree. Reductions should be deterministic; assuming the same number of MPI processes, the contributions to the reduction should always be added in the same order. If you change the number of MPI processes, then I agree that the result of reductions can change, as the order of the sum over points will change, and floating point addition is not associative.
MPI assumes that reduction operations are associative. So does OpenMP, and Carpet. So does SimFactory it its option settings. So does Kranc when it generates code...
-erik
On 22 Feb 2013, at 18:13, Erik Schnetter schnetter@cct.lsu.edu wrote:
On Fri, Feb 22, 2013 at 12:01 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 22 Feb 2013, at 17:43, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi Kentaro,
On Fri, Feb 22, 2013 at 04:48:48PM +0100, Kentaro Takami wrote:
I encountered the problem in which the simulation result is randomly different at each run( of course I used same executable, par file, HPC and so on.).
I see the same for some of my runs, but I wouldn't expect otherwise.
In my case the simulation depends on 'global' quantities, e.g., quantities obtained by reductions. Reductions cannot be done exactly, at least not efficiently. Thus, values obtained via reductions (can) have always an error that is different even when running the same simulation twice, at least when done in parallel. If your simulation depends on this, then these (small) differences can quickly grow, especially in iterative schemes, e.g., hydro.
I disagree. Reductions should be deterministic; assuming the same number of MPI processes, the contributions to the reduction should always be added in the same order. If you change the number of MPI processes, then I agree that the result of reductions can change, as the order of the sum over points will change, and floating point addition is not associative.
MPI assumes that reduction operations are associative.
Ouch. You are right:
http://www.mpi-forum.org/docs/mpi-11-html/node77.html
The operation op is always assumed to be associative. All predefined operations are also assumed to be commutative. Users may define operations that are assumed to be associative, but not commutative. The ``canonical'' evaluation order of a reduction is determined by the ranks of the processes in the group. However, the implementation can take advantage of associativity, or associativity and commutativity in order to change the order of evaluation. This may change the result of the reduction for operations that are not strictly associative and commutative, such as floating point addition.
[] Advice to implementors.
It is strongly recommended that MPI_REDUCE be implemented so that the same result be obtained whenever the function is applied on the same arguments, appearing in the same order. Note that this may prevent optimizations that take advantage of the physical location of processors. ( End of advice to implementors.) The datatype argument of MPI_REDUCE must be compatible with op. Predefined operators work only with the MPI types listed in Sec. Predefined reduce operations and Sec. MINLOC and MAXLOC . User-defined operators may operate on general, derived datatypes. In this case, each argument that the reduce operation is applied to is one element described by such a datatype, which may contain several basic values. This is further explained in Section User-Defined Operations .
So MPI implementors are strongly recommended to ensure that the result is deterministic, but the standard allows them to deviate from this.
So does OpenMP, and Carpet. So does SimFactory it its option settings. So does Kranc when it generates code…
Kranc assumes associativity because it simplifies expressions. This is no different to what the compiler does when it optimises. So if I change the simplification settings in Kranc, or possibly use a newer version of Mathematica, it may generate different code, just like using a different compiler might. But at least with the same original source files and tools, I should be able to do the same experiment more than once and get the same results. If the *runtime* results are allowed to differ, things become very hard to reason about.
On Fri, Feb 22, 2013 at 06:01:19PM +0100, Ian Hinder wrote:
I disagree. Reductions should be deterministic;
That would be nice, but in practice this is not the case.
assuming the same number of MPI processes, the contributions to the reduction should always be added in the same order.
I agree that you should be able to do that, even efficiently.
However, reductions might locally be done faster with OpenMP, and there you loose unless you are very careful to avoid undefined ordering. I don't think Carpet does this, or how difficult that would actually be.
Frank
Hi, Ian,
What quantity is different, and by how much? What happens if you compile without optimisation? Are the differences present in the initial data, or just after some evolution steps? How many steps? Where in the grid do the differences appear?
I compared HDF5 checkpoint files. The difference is only last a few digit, but it is grow up after several iterations.
The checkpoint files in iteration "0" are exactly same and evolution step create such a difference. For my test case, it=1, 2nd step in RK : kxx and so on become random. it=1, 3rd step in RK: hydro variable also become random.
Furthermore , if I use O1 in optimization flag, I don't see difference. On the other hand, if I use O2, random difference is appeared.
Kentaro
Hi,
Finally, I concluded that the comment ( http://software.intel.com/sites/default/files/article/164389/fp-consistency-... ) which Ian suggested, is right, that is, "the random results is caused by variations in the starting address and alignment of the global stack".
So, in order to avoid this problems, I rewrote from "apply_dissipation.F77" to "apply_dissipation.F90" using Fortran 90 style. (I'm attaching the modified files.) Then, we can avoid random results at least my test environments.
If you don't have any objection, could you commit these change to repository?
Kentaro TAKAMI
Kentaro
Thanks for rewriting the code in Fortran 90! The Fortran 77 parts of the Einstein Toolkit should slowly be converted to Fortran 90 anyway.
I notice that you made several changes to the code when converting: (1) You introduce a local array var, which is a copy of the input variable (2) You rewrote the do loops with array expressions
Both are not good for performance. The former is certainly slowing things down, the second complicates an OpenMP parallelisation and makes the code more difficult to read. I am therefore hesitant to apply these. Did you try making only one of these changes, to see whether this would suffice?
The document to which you pointed contains also the suggestion to use the option -fp-model-precise. Did you try this?
As a side remark, in Fortran 90 you can also use a select case statement instead of if statements to choose the dissipation order.
-erik
On Mon, Mar 4, 2013 at 4:35 PM, Kentaro Takami kentaro.takami@aei.mpg.dewrote:
Hi,
Finally, I concluded that the comment ( http://software.intel.com/sites/default/files/article/164389/fp-consistency-... ) which Ian suggested, is right, that is, "the random results is caused by variations in the starting address and alignment of the global stack".
So, in order to avoid this problems, I rewrote from "apply_dissipation.F77" to "apply_dissipation.F90" using Fortran 90 style. (I'm attaching the modified files.) Then, we can avoid random results at least my test environments.
If you don't have any objection, could you commit these change to repository?
Kentaro TAKAMI
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Hi, Erik,
I notice that you made several changes to the code when converting: (1) You introduce a local array var, which is a copy of the input variable (2) You rewrote the do loops with array expressions
Both are not good for performance. The former is certainly slowing things down, the second complicates an OpenMP parallelisation and makes the code more difficult to read. I am therefore hesitant to apply these. Did you try making only one of these changes, to see whether this would suffice?
Maybe we need the changing point (1) to avoid random results, although the copy of array require additional cost (but this array copy is not so expensive compared with the do loop copy.
For changing point (2), we can choose both array expression and do loop style. Unless using OpenMP, array expression style is faster than do loop style, because there is no cost of do loop over head, and compiler can optimize efficiently. However when we use OpenMP, I'm not clear which style is more efficient. The operations in dissipation equation are cheap and simple (so maybe the calculation efficiency is limited by data transfer from cache memory), while OpenMP require large overhead costs of OMP parallelization. Therefore I choose array expression style.
The document to which you pointed contains also the suggestion to use the option -fp-model-precise. Did you try this?
Yes, I tried to use this option. Then we can avoid random results even if we use original f77 code.
As a side remark, in Fortran 90 you can also use a select case statement instead of if statements to choose the dissipation order.
Oh, yes. We should use "select case", because "select case" is more efficient than "if" in recent compiler.
Kentaro
Kentaro
When applying performance optimisations, one has to be careful not to be guided by one's experience, but rather only by measurements. Things behave often in a quite surprising manner. Do you have performance data to support your statements that array expressions are faster than do loops (without OpenMP)? Do you also have performance data to support your statement that OpenMP is slowing things down in this case? With performance data I refer e.g. to Cactus timer output for dissipation for a reasonable setup, running on a reliable machine (e.g. an HPC node).
My personal hypothesis would be that one of the changes you introduce (remove do loops, remove OpenMP parallelization) inhibits some compiler optimisation and thus leads to more consistent results.
As Frank says, reduction operations are not an issue here. I don't see which compiler optimisations would lead to random differences.
Do you have two versions of the produced executable, one that produces random changes and one that doesn't? Could you make these available to me? I would like to compare the produced machine code to see the difference.
-erik
On Tue, Mar 5, 2013 at 3:33 AM, Kentaro Takami kentaro.takami@aei.mpg.dewrote:
Hi, Erik,
I notice that you made several changes to the code when converting: (1) You introduce a local array var, which is a copy of the input
variable
(2) You rewrote the do loops with array expressions
Both are not good for performance. The former is certainly slowing things down, the second complicates an OpenMP parallelisation and makes the code more difficult to read. I am therefore hesitant to apply these. Did you
try
making only one of these changes, to see whether this would suffice?
Maybe we need the changing point (1) to avoid random results, although the copy of array require additional cost (but this array copy is not so expensive compared with the do loop copy.
For changing point (2), we can choose both array expression and do loop style. Unless using OpenMP, array expression style is faster than do loop style, because there is no cost of do loop over head, and compiler can optimize efficiently. However when we use OpenMP, I'm not clear which style is more efficient. The operations in dissipation equation are cheap and simple (so maybe the calculation efficiency is limited by data transfer from cache memory), while OpenMP require large overhead costs of OMP parallelization. Therefore I choose array expression style.
The document to which you pointed contains also the suggestion to use the option -fp-model-precise. Did you try this?
Yes, I tried to use this option. Then we can avoid random results even if we use original f77 code.
As a side remark, in Fortran 90 you can also use a select case statement instead of if statements to choose the dissipation order.
Oh, yes. We should use "select case", because "select case" is more efficient than "if" in recent compiler.
Kentaro
Hi Kentaro,
Just to clarify regarding the OpenMP parallelization. Even though it's supposed to be true that Fortran90 array constructs can be OpenMP parallelized using the !$OMP WORKSHARE directive, it turns out that some compilers (most notably the intel compilers) doesn't actually parallelize these constructs. Instead one processor does all the work. For that unfortunate reason the loop construct is much more efficient than the array construct on many machines.
Cheers,
Peter
On Tue, 5 Mar 2013, Erik Schnetter wrote:
Kentaro When applying performance optimisations, one has to be careful not to be guided by one's experience, but rather only by measurements. Things behave often in a quite surprising manner. Do you have performance data to support your statements that array expressions are faster than do loops (without OpenMP)? Do you also have performance data to support your statement that OpenMP is slowing things down in this case? With performance data I refer e.g. to Cactus timer output for dissipation for a reasonable setup, running on a reliable machine (e.g. an HPC node).
My personal hypothesis would be that one of the changes you introduce (remove do loops, remove OpenMP parallelization) inhibits some compiler optimisation and thus leads to more consistent results.
As Frank says, reduction operations are not an issue here. I don't see which compiler optimisations would lead to random differences.
Do you have two versions of the produced executable, one that produces random changes and one that doesn't? Could you make these available to me? I would like to compare the produced machine code to see the difference.
-erik
On Tue, Mar 5, 2013 at 3:33 AM, Kentaro Takami kentaro.takami@aei.mpg.de wrote: Hi, Erik,
> I notice that you made several changes to the code when converting: > (1) You introduce a local array var, which is a copy of the input variable > (2) You rewrote the do loops with array expressions > > Both are not good for performance. The former is certainly slowing things > down, the second complicates an OpenMP parallelisation and makes the code > more difficult to read. I am therefore hesitant to apply these. Did you try > making only one of these changes, to see whether this would suffice?Maybe we need the changing point (1) to avoid random results, although the copy of array require additional cost (but this array copy is not so expensive compared with the do loop copy.
For changing point (2), we can choose both array expression and do loop style. Unless using OpenMP, array expression style is faster than do loop style, because there is no cost of do loop over head, and compiler can optimize efficiently. However when we use OpenMP, I'm not clear which style is more efficient. The operations in dissipation equation are cheap and simple (so maybe the calculation efficiency is limited by data transfer from cache memory), while OpenMP require large overhead costs of OMP parallelization. Therefore I choose array expression style.
The document to which you pointed contains also the suggestion to
use the
option -fp-model-precise. Did you try this?
Yes, I tried to use this option. Then we can avoid random results even if we use original f77 code.
As a side remark, in Fortran 90 you can also use a select case
statement
instead of if statements to choose the dissipation order.
Oh, yes. We should use "select case", because "select case" is more efficient than "if" in recent compiler.
Kentaro
-- Erik Schnetter schnetter@cct.lsu.edu http://www.perimeterinstitute.ca/personal/eschnetter/
Hi, all,
To Frank,
However using "-fp-model precise" in option list( i.e., this option is used by all fortran77 code), the run will become much slow down.
Out of interest: how much is "much"?
I didn't test it. However the document( p11 in http://software.intel.com/sites/default/files/article/164389/fp-consistency-... ) say that the performance is reduced by 12-15%.
To Peter,
Just to clarify regarding the OpenMP parallelization. Even though it's supposed to be true that Fortran90 array constructs can be OpenMP parallelized using the !$OMP WORKSHARE directive, it turns out that some compilers (most notably the intel compilers) doesn't actually parallelize these constructs. Instead one processor does all the work. For that unfortunate reason the loop construct is much more efficient than the array construct on many machines.
Thank you for telling me about it. I have never used "!$OMP WORKSHARE" directive. So I tested this directive in the following performance test.
To Erik,
Do you have performance data to support your statements that array expressions are faster than do loops (without OpenMP)?
Intel compiler document such as http://www-h.eng.cam.ac.uk/help/languages/fortran/intelfortran/for_ug2.pdf , say "Rather than use explicit loops for array access, use elemental array operations...(P20)". So I thought that array expressions are faster than do-loops, although I had never compared these.
I'm attaching data of performance test for attached par file. Due to these data, both expressions are almost same.
Do you also have performance data to support your statement that OpenMP is slowing things down in this case?
Please see the attached file. Apparently the case (4) is the fastest, that is, OMP palatalization shouldn't be used for these small section.
My personal hypothesis would be that one of the changes you introduce (remove do loops, remove OpenMP parallelization) inhibits some compiler optimisation and thus leads to more consistent results.
Already I mentioned, "introducing a local array var, which is a copy of the input variable" prevent from random results.
Do you have two versions of the produced executable, one that produces random changes and one that doesn't? Could you make these available to me? I would like to compare the produced machine code to see the difference.
I guess you can access damiana@AEI. You can use the following executable files in "/lustre/datura/takami/Share_Dir/EXE__Test_Dissipation/".
For random results, "cactus_GRHydro.F77--datura.OMP-NO_PRECISE-NO".
For consistent results, "cactus_GRHydro.F77--datura.OMP-NO_PRECISE-YES".
Here, both are exactly same source programs and compile options except for "-fp-model precise" option.
In order to compare the results, usually I use the following command: $h5diff \ OUTPUT_DATA_1/CHECKPOINT/checkpoint.chkpt.it_5.file_0.h5 \ OUTPUT_DATA_2/CHECKPOINT/checkpoint.chkpt.it_5.file_0.h5
Kentaro TAKAMI
Hi,
On Wed, Mar 06, 2013 at 06:45:17PM +0100, Kentaro Takami wrote:
http://software.intel.com/sites/default/files/article/164389/fp-consistency-... say that the performance is reduced by 12-15%.
I would be careful with using such numbers. I would expect this to depend quite heavily on your specific application. (For these numbers Intel used the Speccpu2006fp benchmark.) I wonder what the impact on one of our typical simulations is.
In the end it's a trade-off between accuracy and speed.
Frank
On Wed, Mar 6, 2013 at 12:45 PM, Kentaro Takami kentaro.takami@aei.mpg.dewrote:
Do you have two versions of the produced executable, one that produces
random changes and one that doesn't? Could you make these available to
me? I
would like to compare the produced machine code to see the difference.
I guess you can access damiana@AEI. You can use the following executable files in "/lustre/datura/takami/Share_Dir/EXE__Test_Dissipation/".
For random results, "cactus_GRHydro.F77--datura.OMP-NO_PRECISE-NO".
For consistent results, "cactus_GRHydro.F77--datura.OMP-NO_PRECISE-YES".
Here, both are exactly same source programs and compile options except for "-fp-model precise" option.
Kentaro
Yes, I have access to Damiana.
These two executables are not useful to me, since you said you are not interested in using the -fp-model-precise option anyway. I therefore see no reason to examine it. I am rather interested in comparing the executables for (a) the original version and (b) a version that you "like".
-erik
Hi, Erik,
These two executables are not useful to me, since you said you are not interested in using the -fp-model-precise option anyway. I therefore see no reason to examine it. I am rather interested in comparing the executables for (a) the original version and (b) a version that you "like".
OK. In same directory, you can find all executables used by performance test. In order to find the reason why random changes are produced or not, "cactus_GRHydro.F90_DoStyle--datura.OMP-NO_PRECISE-NO" is reasonable, because "do-loop" style is used. I note only "Dissipation::order = 5" can be used in this executable.
Do you also have performance data to support your statement that OpenMP is slowing things down in this case?
Please see the attached file. Apparently the case (4) is the fastest, that is, OMP palatalization shouldn't be used for these small section.
If I read the numbers correctly, then the original code is the fastest in all cases, sometimes by a factor of five.
Oh, you are right. Sorry, I was seeing consuming time, i.e., "getrusage".
Due to the data, this compiler doesn't parallelize "!$OMP PARALLEL WORKSHARE" properly. So, as Peter pointed out, many compiler still can't use this directive, and it is safe to use "!$OMP PARALLEL DO".
OK, I agree that we use "do-loop" + "OpenMP" now.
Forthermore, "introducing a local array var" to avoid random change, derive slow down, especially for OpenMP.
Kentaro TAKAMI
On Wed, Mar 6, 2013 at 12:45 PM, Kentaro Takami kentaro.takami@aei.mpg.dewrote:
Do you also have performance data to support your statement that
OpenMP is slowing things down in this case?
Please see the attached file. Apparently the case (4) is the fastest, that is, OMP palatalization shouldn't be used for these small section.
If I read the numbers correctly, then the original code is the fastest in all cases, sometimes by a factor of five.
-erik
On Mon, Mar 04, 2013 at 10:35:32PM +0100, Kentaro Takami wrote:
If you don't have any objection, could you commit these change to repository?
Thanks for taking your time looking into this issue. Using Fortran 90 makes the code looks nicer indeed. However, I hesitate to apply these changes.
First of all, the document talks about reductions. This code doesn't do any reduction. Each thread writes only it's own little piece of memory. Changes in order of thread execution should not change anything here. Also, inside the loop no external functions are called that could influence the outcome. I am curious to see which of your (many) changes is the real workaround for the problem. Is it the initial copy of the input array (which is probably quite expensive)?
You remove the openmp parallelization. Why?
Did you try the suggested compiler options (-fp-model precise)?
Frank
Hi, Frank,
First of all, the document talks about reductions. This code doesn't do any reduction.
No, I'm talking about "Second Example from WRF" in the document. I thought this example is related with our case. The document also say that Intel 11 compiler avoid this problem, but I'm still seeing this random result in Intel 11 compiler. (So our situation might be a little bit different from the example.)
You remove the openmp parallelization. Why?
I explained this reason in previous message.
Did you try the suggested compiler options (-fp-model precise)?
Yes, I did. Then we can avoid random result. However using "-fp-model precise" in option list( i.e., this option is used by all fortran77 code), the run will become much slow down.
Kentaro TAKAMI
First of all, I should give you more information.
In my test, I don't use OpenMP. So I use only MPI. I put dissipation only for spacetime variable such as gij, kij and so on. I check these difference for GF in checkpoint files. The difference appear from 2nd step of RK at 1st iteration, where the difference are for last a few digit, but grow up as increasing iterations.
In the following case we don't see difference. *don't activate Dissipation thorn. *using only 1 refinement level. *using certain special dx=dy=dz.
Further strange behavior: When I insert non-essential line such as "WRITE(*,*)" statement in "apply_dissipation.F77", sometimes the difference is disappeared. So, I doubted incorrect pointer position.
Hi, Frank,
Finally I'm doubting the origin of the problem is "CactusNumerical/Dissipation", because when I don't activate this thorn, the problem can be disappeared.
This is interesting. So, you say you don't see random changes when Dissipation isn't active, but you do when you activate it, and everything else is the same? Which variables do you apply dissipation to?
Yes, if we don't use Dissipation, HDF5 checkpoint files can become exactly same, at least my test case. I added dissipation for the following variables: ML_BSSN::ML_log_confac ML_BSSN::ML_metric ML_BSSN::ML_trace_curv ML_BSSN::ML_curv ML_BSSN::ML_Gamma ML_BSSN::ML_lapse ML_BSSN::ML_shift ML_BSSN::ML_dtlapse ML_BSSN::ML_dtshift
Kentaro
On 22 Feb 2013, at 18:38, Kentaro Takami kentaro.takami@aei.mpg.de wrote:
First of all, I should give you more information.
In my test, I don't use OpenMP. So I use only MPI. I put dissipation only for spacetime variable such as gij, kij and so on. I check these difference for GF in checkpoint files. The difference appear from 2nd step of RK at 1st iteration, where the difference are for last a few digit, but grow up as increasing iterations.
In the following case we don't see difference. *don't activate Dissipation thorn. *using only 1 refinement level. *using certain special dx=dy=dz.
Further strange behavior: When I insert non-essential line such as "WRITE(*,*)" statement in "apply_dissipation.F77", sometimes the difference is disappeared. So, I doubted incorrect pointer position.
Are you using sufficient ghost zones?
Hi, Frank,
Finally I'm doubting the origin of the problem is "CactusNumerical/Dissipation", because when I don't activate this thorn, the problem can be disappeared.
This is interesting. So, you say you don't see random changes when Dissipation isn't active, but you do when you activate it, and everything else is the same? Which variables do you apply dissipation to?
Yes, if we don't use Dissipation, HDF5 checkpoint files can become exactly same, at least my test case. I added dissipation for the following variables: ML_BSSN::ML_log_confac ML_BSSN::ML_metric ML_BSSN::ML_trace_curv ML_BSSN::ML_curv ML_BSSN::ML_Gamma ML_BSSN::ML_lapse ML_BSSN::ML_shift ML_BSSN::ML_dtlapse ML_BSSN::ML_dtshift
Kentaro _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On Fri, Feb 22, 2013 at 06:49:46PM +0100, Ian Hinder wrote:
Are you using sufficient ghost zones?
Dissipation checks for that and aborts if it is not the case. The only reason I can think of right now why random things could happen is if the primary variables (not the rhs) are not updated in the ghost- or boundary zones before Dissipation is called in post_rhs.
Frank
On Feb 22, 2013, at 10:38 AM, Kentaro Takami wrote:
Further strange behavior: When I insert non-essential line such as "WRITE(*,*)" statement in "apply_dissipation.F77", sometimes the difference is disappeared. So, I doubted incorrect pointer position.
could it be related to compiler optimization?
I remember seeing things similar to this in the past, when adding a line like that was changing the behavior of the code. In my case, it was having or not a segfault, which at the end depended on how the compiler was actually compiling the code and that seemed to be somehow dependent on the length of the routine. This is a problem I had few years ago so I don't remember exactly all the details, but we fixed it by changing the code in order to be sure that certain operations were always executed in the same order (I think it had to do with an if condition being located inside or outside a do loop).
Cheers, Bruno
Dr. Bruno Giacomazzo JILA - University of Colorado 440 UCB Boulder, CO 80309 USA
Tel. : +1 303-492-5170 Fax : +1 303-492-5235 email : bruno.giacomazzo@jila.colorado.edu web: http://www.brunogiacomazzo.org
---------------------------------------------------------------------- There are only 10 types of people in the world: Those who understand binary, and those who don't ----------------------------------------------------------------------
users@lists.einsteintoolkit.org