Hi,
I just checked out a new copy of the Einstein Toolkit and compiled, and got this (and a lot of similar messages):
An internal threshold was exceeded for routine _Z20ML_BSSN_O2_RHS2_BodyPK4_cGHiiPKdS3_S3_PKiS5_iPrKPd and optimization level may be reduced. See http://software.intel.com/en-us/articles/internal-threshold-was-exceeded for more information and advice.
In addition, McLachlan seems to compile for a very long time and uses a lot of memory (several GB per compiler process), for the Intel compiler.
Did I miss this in the past or did something change in McLachlan which could trigger this?
Frank
The compiler options for your machine may have changed.
-erik
On Mon, Sep 26, 2011 at 10:08 PM, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi,
I just checked out a new copy of the Einstein Toolkit and compiled, and got this (and a lot of similar messages):
An internal threshold was exceeded for routine _Z20ML_BSSN_O2_RHS2_BodyPK4_cGHiiPKdS3_S3_PKiS5_iPrKPd and optimization level may be reduced. See http://software.intel.com/en-us/articles/internal-threshold-was-exceeded for more information and advice.
In addition, McLachlan seems to compile for a very long time and uses a lot of memory (several GB per compiler process), for the Intel compiler.
Did I miss this in the past or did something change in McLachlan which could trigger this?
Frank
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Hi
On Mon, Sep 26, 2011 at 09:08:40PM -0500, Frank Loeffler wrote:
In addition, McLachlan seems to compile for a very long time and uses a lot of memory (several GB per compiler process), for the Intel compiler.
Actually, it makes it unusable: even one compiler process now eats above 10GB of Ram, on a 8GB machine. Am I the only one with this problem? This is on my workstation (numrel07) where this used to work for quite a while.
Frank
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
On 27 Sep 2011, at 06:04, Frank Loeffler wrote:
Hi
On Mon, Sep 26, 2011 at 09:08:40PM -0500, Frank Loeffler wrote:
In addition, McLachlan seems to compile for a very long time and uses a lot of memory (several GB per compiler process), for the Intel compiler.
Actually, it makes it unusable: even one compiler process now eats above 10GB of Ram, on a 8GB machine. Am I the only one with this problem? This is on my workstation (numrel07) where this used to work for quite a while.
Yes, McLachlan has changed. Kranc can now select the finite difference operator based on a run-time parameter. This eliminates the need for multiple versions of the McLachlan thorn. You can just use ML_BSSN and set fdOrder to 2, 4, 6 or 8. For compatibility purposes, for one release, we keep the existing thorns and set the default of fdOrder to the corresponding thorn order.
All finite difference methods are compiled in, and selected with a switch statement in the inner loop. This makes the inner loop much larger than it was before, which must be causing problems for the compiler (even though it should be able to tell that the separate branches of the switch statement are independent and can never be optimised together). This should not have a run-time performance impact because the additional instructions are not executed, and Erik tested that the generated code didn't look any worse than before. I also saw negligible changes in speed.
The reason that we didn't see any problems and you did is that we were using
VECTORISE_INLINE = no
as in datura.cfg, whereas the default is "yes", and the numrel.cfg file leaves it set to the default. Leaving this set to "yes" means that the Vectors thorn will attempt to inline vectorised operations, and the Vectors thorn also uses this to tell Kranc to use inlining for the finite differencing operators rather than functions. Using functions for the derivatives allows the code to fit into the instruction cache in many cases where it previously would not, for example when using 8th order with multipatch. It is "yes" by default as one would expect function calls to be slower than inlining when instruction-cache misses are not relevant.
I can reproduce your problem. I created a little wrapper for icpc:
icpc-time: /usr/bin/time -f 'WALLTIME=%E s, MAXRSS=%M kB' /cluster/Compiler/Intel/11.1.072/bin/intel64/icpc "${@}"
and changed the optionlist to use this instead of icpc. I then compiled ML_BSSN_O2 using this new optionlist. With VECTORISE_INLINE = no (datura's default), I get
COMPILING /home/ianhin/Cactus/EinsteinToolkit/arrangements/McLachlan/ML_BSSN_O2/src/ML_BSSN_O2_RHS2.cc WALLTIME=0:10.61 s, MAXRSS=811584 kB
i.e. it uses 800 MB to compile ML_BSSN_O2_RHS2.cc in 10 seconds.
I then set VECTORISE_INLINE = yes (the default), and ML_BSSN_O2_RHS2 printed the threshold warning that you saw and started taking many GB of RAM and didn't finish compiling after about 10 minutes.
So a fix to your problem is to set VECTORISE_INLINE = no in your optionlist and reconfigure. Alternatively, you could temporarily remove the 4th, 6th and 8th order code from ML_BSSN_O2 by editing McLachlan/m/McLachlan_BSSN.m and modifying intParameters/fdOrder/AllowedValues to be {derivOrder} instead of {2,4,6,8}. Then type "make McLachlan_BSSN.out" in the m directory to regenerate the thorn. Then rebuild.
Erik: should we make VECTORISE_INLINE = no the default in Vectors?
- -- Ian Hinder ian.hinder@aei.mpg.de
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
On 27 Sep 2011, at 14:20, Ian Hinder wrote:
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
On 27 Sep 2011, at 06:04, Frank Loeffler wrote:
Hi
On Mon, Sep 26, 2011 at 09:08:40PM -0500, Frank Loeffler wrote:
In addition, McLachlan seems to compile for a very long time and uses a lot of memory (several GB per compiler process), for the Intel compiler.
Actually, it makes it unusable: even one compiler process now eats above 10GB of Ram, on a 8GB machine. Am I the only one with this problem? This is on my workstation (numrel07) where this used to work for quite a while.
Yes, McLachlan has changed. Kranc can now select the finite difference operator based on a run-time parameter. This eliminates the need for multiple versions of the McLachlan thorn. You can just use ML_BSSN and set fdOrder to 2, 4, 6 or 8. For compatibility purposes, for one release, we keep the existing thorns and set the default of fdOrder to the corresponding thorn order.
All finite difference methods are compiled in, and selected with a switch statement in the inner loop. This makes the inner loop much larger than it was before, which must be causing problems for the compiler (even though it should be able to tell that the separate branches of the switch statement are independent and can never be optimised together).
I just checked, and even before combining all the FD orders into one thorn, the compilation with VECTORISE_INLINE = "yes" was slow and used a lot of memory (> 8 GB) for 8th order and multipatch with vectorisation. The compiler memory usage does seem to have increased by combining the FD orders however. Given that production binary black hole simulations typically use 8th order these days, I think we should just use functions for the difference operators by setting VECTORISE_INLINE = no.
- -- Ian Hinder ian.hinder@aei.mpg.de
Hi,
On Tue, Sep 27, 2011 at 03:02:42PM +0200, Ian Hinder wrote:
I just checked, and even before combining all the FD orders into one thorn, the compilation with VECTORISE_INLINE = "yes" was slow and used a lot of memory (> 8 GB) for 8th order and multipatch with vectorisation. The compiler memory usage does seem to have increased by combining the FD orders however. Given that production binary black hole simulations typically use 8th order these days, I think we should just use functions for the difference operators by setting VECTORISE_INLINE = no.
Can we set this unconditionally for McLachlan only? Otherwise people are bound to stumble over this when they try the released version.
Frank
Setting this unconditionally doesn't make sense; there may be another compiler, or another version, or a different set of compiler options which doesn't have such a dramatic slowdown, and which produces much slower code in this case. Things need to remain configurable.
Let's fix this particular problem first, and then look for a general solution once we know the behaviour on all the other machines.
-erik
On Tue, Sep 27, 2011 at 10:56 AM, Frank Loeffler knarf@cct.lsu.edu wrote:
Hi,
On Tue, Sep 27, 2011 at 03:02:42PM +0200, Ian Hinder wrote:
I just checked, and even before combining all the FD orders into one thorn, the compilation with VECTORISE_INLINE = "yes" was slow and used a lot of memory (> 8 GB) for 8th order and multipatch with vectorisation. The compiler memory usage does seem to have increased by combining the FD orders however. Given that production binary black hole simulations typically use 8th order these days, I think we should just use functions for the difference operators by setting VECTORISE_INLINE = no.
Can we set this unconditionally for McLachlan only? Otherwise people are bound to stumble over this when they try the released version.
Frank
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
On 27 Sep 2011, at 16:56, Frank Loeffler wrote:
Hi,
On Tue, Sep 27, 2011 at 03:02:42PM +0200, Ian Hinder wrote:
I just checked, and even before combining all the FD orders into one thorn, the compilation with VECTORISE_INLINE = "yes" was slow and used a lot of memory (> 8 GB) for 8th order and multipatch with vectorisation. The compiler memory usage does seem to have increased by combining the FD orders however. Given that production binary black hole simulations typically use 8th order these days, I think we should just use functions for the difference operators by setting VECTORISE_INLINE = no.
Can we set this unconditionally for McLachlan only? Otherwise people are bound to stumble over this when they try the released version.
What other thorn would this affect than McLachlan?
- -- Ian Hinder ian.hinder@aei.mpg.de
WeylScal4, the WaveToy/ADM examples (not important), and the codes that other users create with Kranc.
-erik
On Tue, Sep 27, 2011 at 11:06 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
On 27 Sep 2011, at 16:56, Frank Loeffler wrote:
Hi,
On Tue, Sep 27, 2011 at 03:02:42PM +0200, Ian Hinder wrote:
I just checked, and even before combining all the FD orders into one thorn, the compilation with VECTORISE_INLINE = "yes" was slow and used a lot of memory (> 8 GB) for 8th order and multipatch with vectorisation. The compiler memory usage does seem to have increased by combining the FD orders however. Given that production binary black hole simulations typically use 8th order these days, I think we should just use functions for the difference operators by setting VECTORISE_INLINE = no.
Can we set this unconditionally for McLachlan only? Otherwise people are bound to stumble over this when they try the released version.
What other thorn would this affect than McLachlan?
Ian Hinder ian.hinder@aei.mpg.de
-----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.8 (Darwin)
iEYEARECAAYFAk6B5ncACgkQF1LN8Zj+CehDuwCfYThPELUoj9v66mProAOq/HOF tcsAn3ZVfzu/uzp5FaPvmefviV4Un4Rs =KYht -----END PGP SIGNATURE----- _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 27 Sep 2011, at 17:16, Erik Schnetter wrote:
WeylScal4, the WaveToy/ADM examples (not important), and the codes that other users create with Kranc.
And why would changing the default for VECTORISE_INLINE adversely affect these? I'm trying to understand Frank's statement that "people are bound to stumble over this when they try the released version". Frank, did you mean the soon-to-be (ET_2011_10) released version? Or the previously released version? If we change the default to VECTORISE_INLINE = no now, then no-one will run into this problem in the soon-to-be released version. Since the change to McLachlan didn't exist in the previously released version, it won't affect those.
-erik
On Tue, Sep 27, 2011 at 11:06 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
On 27 Sep 2011, at 16:56, Frank Loeffler wrote:
Hi,
On Tue, Sep 27, 2011 at 03:02:42PM +0200, Ian Hinder wrote:
I just checked, and even before combining all the FD orders into one thorn, the compilation with VECTORISE_INLINE = "yes" was slow and used a lot of memory (> 8 GB) for 8th order and multipatch with vectorisation. The compiler memory usage does seem to have increased by combining the FD orders however. Given that production binary black hole simulations typically use 8th order these days, I think we should just use functions for the difference operators by setting VECTORISE_INLINE = no.
Can we set this unconditionally for McLachlan only? Otherwise people are bound to stumble over this when they try the released version.
What other thorn would this affect than McLachlan?
Ian Hinder ian.hinder@aei.mpg.de
-----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.8 (Darwin)
iEYEARECAAYFAk6B5ncACgkQF1LN8Zj+CehDuwCfYThPELUoj9v66mProAOq/HOF tcsAn3ZVfzu/uzp5FaPvmefviV4Un4Rs =KYht -----END PGP SIGNATURE----- _______________________________________________ Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
-- Erik Schnetter schnetter@cct.lsu.edu http://www.cct.lsu.edu/~eschnett/
Hello Ian, all,
And why would changing the default for VECTORISE_INLINE adversely affect these? I'm trying to understand Frank's statement that "people are bound to stumble over this when they try the released version". Frank, did you mean the soon-to-be (ET_2011_10) released version? Or the previously released version? If we change the default to VECTORISE_INLINE = no now, then no-one will run into this problem in the soon-to-be released version. Since the change to McLachlan didn't exist in the previously released version, it won't affect those.
Well someone might have non-simfactory options lists (those exist :-) ) and *that* one might say VECTORISE = yes which they want for their own private Kranc generated thorn. This way they would stumble across this.
I had wondered about that before: the method to support multiple orders of differentiation used in WeylScal4 and Kranc2BSSN.m was not used because it requires "unrolling" (exploding really) all choices for the combination of parameters at compile time? With lots of "schedule ... AS" statements the schedule does not look any different.
What happens if a user tries VECTORISE=yes on say a head of Ranger? Will they bring down the node or just have unexplainable compile errors in Kranc's generated code?
Yours, Roland
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
On 27 Sep 2011, at 17:33, Roland Haas wrote:
Hello Ian, all,
And why would changing the default for VECTORISE_INLINE adversely affect these? I'm trying to understand Frank's statement that "people are bound to stumble over this when they try the released version". Frank, did you mean the soon-to-be (ET_2011_10) released version? Or the previously released version? If we change the default to VECTORISE_INLINE = no now, then no-one will run into this problem in the soon-to-be released version. Since the change to McLachlan didn't exist in the previously released version, it won't affect those.
Well someone might have non-simfactory options lists (those exist :-) ) and *that* one might say VECTORISE = yes which they want for their own private Kranc generated thorn. This way they would stumble across this.
If we change the default, then the default will be VECTORISE_INLINE = no. We have to have a default, and it's not clear which is better. Vectorisation is new in this release of the ET, so no one using a previous release should be affected. We know that McLachlan prefers VECTORISE_INLINE = no for high order differencing, otherwise the compiler takes too much memory. So I say we should set the default to the only thing we know is a definite improvement for the thorns in the ET.
I had wondered about that before: the method to support multiple orders of differentiation used in WeylScal4 and Kranc2BSSN.m was not used because it requires "unrolling" (exploding really) all choices for the combination of parameters at compile time? With lots of "schedule ... AS" statements the schedule does not look any different.
? The addition of this feature of Kranc was intended to avoid this proliferation of parameterised calculations. I'm not sure what you are trying to say...
What happens if a user tries VECTORISE=yes on say a head of Ranger? Will they bring down the node or just have unexplainable compile errors in Kranc's generated code?
If they have VECTORISE = yes (the proposed default for the new release of the ET) and VECTORISE_INLINE = yes (the current default, which I propose should be changed to "no"), and they try to compile McLachlan, the behaviour will be as Frank saw, which is that you get a warning and the compiler runs out of memory. If the Ranger head node is configured such that a user can bring down the system by running a program which takes up a lot of memory, then the head node will go down. This is common for default linux installs (I don' t know why), but Ranger has a ulimit setting of 8 GB for virtual memory, so I think the process will be terminated when it uses more than 8 GB, which is a sensible precaution for a production system.
- -- Ian Hinder ian.hinder@aei.mpg.de
On Tue, Sep 27, 2011 at 11:50 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
If they have VECTORISE = yes (the proposed default for the new release of the ET) and VECTORISE_INLINE = yes (the current default, which I propose should be changed to "no"), and they try to compile McLachlan, the behaviour will be as Frank saw, which is that you get a warning and the compiler runs out of memory. If the Ranger head node is configured such that a user can bring down the system by running a program which takes up a lot of memory, then the head node will go down. This is common for default linux installs (I don' t know why), but Ranger has a ulimit setting of 8 GB for virtual memory, so I think the process will be terminated when it uses more than 8 GB, which is a sensible precaution for a production system.
This depends very much on the compiler, the compiler version, and the options chosen. A good compiler won't use too much memory when optimizing; this is really a compiler bug (that we should report).
I regularly encounter cases where a compiler takes a long time, runs out of memory, or outputs an error message that it stopped optimizing. Or the compiler crashes, sometimes in debug mode and without optimizing. What I usually do is either modify the source or the Simfactory options. We simply need to do the same here. (There's no need to be afraid that Ranger will explode, it didn't explode in the past either.)
Frank, which machine was this? Which compiler version? What options did you use? Could you open a trac issue about this?
-erik
On 27 Sep 2011, at 18:00, Erik Schnetter wrote:
On Tue, Sep 27, 2011 at 11:50 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
If they have VECTORISE = yes (the proposed default for the new release of the ET) and VECTORISE_INLINE = yes (the current default, which I propose should be changed to "no"), and they try to compile McLachlan, the behaviour will be as Frank saw, which is that you get a warning and the compiler runs out of memory. If the Ranger head node is configured such that a user can bring down the system by running a program which takes up a lot of memory, then the head node will go down. This is common for default linux installs (I don' t know why), but Ranger has a ulimit setting of 8 GB for virtual memory, so I think the process will be terminated when it uses more than 8 GB, which is a sensible precaution for a production system.
This depends very much on the compiler, the compiler version, and the options chosen. A good compiler won't use too much memory when optimizing; this is really a compiler bug (that we should report).
I regularly encounter cases where a compiler takes a long time, runs out of memory, or outputs an error message that it stopped optimizing. Or the compiler crashes, sometimes in debug mode and without optimizing. What I usually do is either modify the source or the Simfactory options. We simply need to do the same here. (There's no need to be afraid that Ranger will explode, it didn't explode in the past either.)
Frank, which machine was this? Which compiler version? What options did you use? Could you open a trac issue about this?
This was numrel07, he was using numrel.cfg from simfactory. I reproduced his problem in Datura by setting VECTORISE_INLINE = yes. I think we should just change the default in Vectors to VECTORISE_INLINE = no. Do you think this is a bad idea? If so, why?
I created the VECTORISE entries in the Simfactory MDB files so that the generated code would be most efficient. There is in particular a difference between Intel and AMD systems, having to do with the size of the instruction cache.
Yes, we can change the default, but please keep the behaviour of the current machines the same unless there is a problem when compiling. Just switching off inlining everywhere would undo my work, even in cases where this is not necessary. In other words, when changing the default, I ask you do add VECTORISE_INLINE = yes to all MDB entries that don't set anything at the moment.
-erik
On Tue, Sep 27, 2011 at 12:07 PM, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 27 Sep 2011, at 18:00, Erik Schnetter wrote:
On Tue, Sep 27, 2011 at 11:50 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
If they have VECTORISE = yes (the proposed default for the new release of the ET) and VECTORISE_INLINE = yes (the current default, which I propose should be changed to "no"), and they try to compile McLachlan, the behaviour will be as Frank saw, which is that you get a warning and the compiler runs out of memory. If the Ranger head node is configured such that a user can bring down the system by running a program which takes up a lot of memory, then the head node will go down. This is common for default linux installs (I don' t know why), but Ranger has a ulimit setting of 8 GB for virtual memory, so I think the process will be terminated when it uses more than 8 GB, which is a sensible precaution for a production system.
This depends very much on the compiler, the compiler version, and the options chosen. A good compiler won't use too much memory when optimizing; this is really a compiler bug (that we should report).
I regularly encounter cases where a compiler takes a long time, runs out of memory, or outputs an error message that it stopped optimizing. Or the compiler crashes, sometimes in debug mode and without optimizing. What I usually do is either modify the source or the Simfactory options. We simply need to do the same here. (There's no need to be afraid that Ranger will explode, it didn't explode in the past either.)
Frank, which machine was this? Which compiler version? What options did you use? Could you open a trac issue about this?
This was numrel07, he was using numrel.cfg from simfactory. I reproduced his problem in Datura by setting VECTORISE_INLINE = yes. I think we should just change the default in Vectors to VECTORISE_INLINE = no. Do you think this is a bad idea? If so, why?
-- Ian Hinder ian.hinder@aei.mpg.de
Hi,
On Tue, Sep 27, 2011 at 12:00:16PM -0400, Erik Schnetter wrote:
I regularly encounter cases where a compiler takes a long time, runs out of memory, or outputs an error message that it stopped optimizing.
Right, this is a compiler bug.
Or the compiler crashes, sometimes in debug mode and without optimizing. What I usually do is either modify the source or the Simfactory options.
Would it be possible to change the source code such that it is more compiler friendly and still shows the same efficiency?
Frank
Hello Ian, all,
If we change the default, then the default will be VECTORISE_INLINE = no. We have to have a default, and it's not clear which is better. Vectorisation is new in this release of the ET, so no one using a previous release should be affected. We know that McLachlan prefers VECTORISE_INLINE = no for high order differencing, otherwise the compiler takes too much memory. So I say we should set the default to the only thing we know is a definite improvement for the thorns in the ET.
Sounds reasonable. A comment why we make this choice in the option lists might be helpful though to avoid users blindly turning it on again since "inlining is always better" [see http://www.mjmwired.net/kernel/Documentation/CodingStyle#692 :-)].
? The addition of this feature of Kranc was intended to avoid this proliferation of parameterised calculations. I'm not sure what you are trying to say...
Ok, this is more or less what I suspected. My point would have been that the (smaller) parameterized calculations might have avoided the out-of-memory problems (though Erik says that 8th order always used lots of memory). Not having the number of routines explode is certainly a nice thing.
know why), but Ranger has a ulimit setting of 8 GB for virtual memory, so I think the process will be terminated when it uses more than 8 GB, which is a sensible precaution for a production system.
ok. Then we will not get lots of angry emails from Ranger users that brought down the head node. Thanks.
Yours, Roland
Hi,
On Tue, Sep 27, 2011 at 05:50:03PM +0200, Ian Hinder wrote:
but Ranger has a ulimit setting of 8 GB for virtual memory, so I think the process will be terminated when it uses more than 8 GB, which is a sensible precaution for a production system.
Yes, but then we typically compile in parallel and there is more than one file which produces this problem...
Frank
Hi,
On Tue, Sep 27, 2011 at 05:23:15PM +0200, Ian Hinder wrote:
And why would changing the default for VECTORISE_INLINE adversely affect these? I'm trying to understand Frank's statement that "people are bound to stumble over this when they try the released version".
I was thinking about the scenario when someone takes the ET, and wants to compile it on their machine. This should at least work, even if not most efficiently.
Frank
On Tue, Sep 27, 2011 at 8:20 AM, Ian Hinder ian.hinder@aei.mpg.de wrote:
Erik: should we make VECTORISE_INLINE = no the default in Vectors?
Yes, if we update all Simfactory entries so that the behaviour on current machines isn't changed.
-erik
users@lists.einsteintoolkit.org