Hello all,
NERSC's edison cluster specifies -xHost in its cfg file. However login nodes and compute nodes are actually different (Sandy Bridge vs. Ivy Bridge cpus).
A quick test removing -xHost reveals a possible reason: without -xHost the compiled exectuables do not run on the head node. No for the actual question: is it worthwhile to claim cross compilation (and the requirement to have to specify endianess, type sizes etc manually) to possibly gain some more speed?
Given that Ivy and Sandy Bridge are tick and tock I would not expect much of a gain (though maybe Ivy Bridge actually offers wider AVX instructions?).
At least I will add a comment to edison.cfg explaining why either -xHost is used or why a cross compilation is required.
Yours, Roland
It's strange that code without "-xHost" does not run on the head node. What other compiler options are there? The only reason I can see is that the Intel compiler has been installed in a special way to use additional compiler flags, and these flags make the code not work any more on Sandy Bridge.
To my knowledge, the processors accept the same machine instructions. The CPU tuning might be different.
Running short tests or compiled utilities on the head nodes is very convenient. I would continue to make sure the code runs everywhere until there is a proven performance benefit.
-erik
On Sun, Sep 18, 2016 at 2:36 PM, Roland Haas rhaas@illinois.edu wrote:
Hello all,
NERSC's edison cluster specifies -xHost in its cfg file. However login nodes and compute nodes are actually different (Sandy Bridge vs. Ivy Bridge cpus).
A quick test removing -xHost reveals a possible reason: without -xHost the compiled exectuables do not run on the head node. No for the actual question: is it worthwhile to claim cross compilation (and the requirement to have to specify endianess, type sizes etc manually) to possibly gain some more speed?
Given that Ivy and Sandy Bridge are tick and tock I would not expect much of a gain (though maybe Ivy Bridge actually offers wider AVX instructions?).
At least I will add a comment to edison.cfg explaining why either -xHost is used or why a cross compilation is required.
Yours, Roland
-- My email is as private as my paper mail. I therefore support encrypting and signing email messages. Get my PGP key from http://keys.gnupg.net.
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
Hello Erik,
It's strange that code without "-xHost" does not run on the head node. What other compiler options are there? The only reason I can see is that the Intel compiler has been installed in a special way to use additional compiler flags, and these flags make the code not work any more on Sandy Bridge.
It is some runtime detection that fails. I get: --8<-- Please verify that both the operating system and the processor support Intel(R) F16C instructions. --8<-- and adding -xHost complains that "icc: command line warning #10121: overriding '-xCORE-AVX-I' with '-xHost'" so it seesm as if some vectorization is enabled by the cray cc wrapper.
To my knowledge, the processors accept the same machine instructions. The CPU tuning might be different.
Is this also true for vectorization? Or maybe this really is just a lack of OS support for the AVX registers on the login nodes. -xAVX seems to work fine and the only difference (according to the man page for icpc) are "Float-16 conversion instructions and the RDRND instruction". Float16 conversion would fit the F16C acronym in the error message.
Running short tests or compiled utilities on the head nodes is very convenient. I would continue to make sure the code runs everywhere until there is a proven performance benefit.
Yes, that is what I would do as well.
Yours, Roland
On Sun, Sep 18, 2016 at 4:24 PM, Roland Haas rhaas@illinois.edu wrote:
Hello Erik,
It's strange that code without "-xHost" does not run on the head node.
What
other compiler options are there? The only reason I can see is that the Intel compiler has been installed in a special way to use additional compiler flags, and these flags make the code not work any more on Sandy Bridge.
It is some runtime detection that fails. I get: --8<-- Please verify that both the operating system and the processor support Intel(R) F16C instructions. --8<-- and adding -xHost complains that "icc: command line warning #10121: overriding '-xCORE-AVX-I' with '-xHost'" so it seesm as if some vectorization is enabled by the cray cc wrapper.
To my knowledge, the processors accept the same machine instructions. The CPU tuning might be different.
Is this also true for vectorization? Or maybe this really is just a lack of OS support for the AVX registers on the login nodes. -xAVX seems to work fine and the only difference (according to the man page for icpc) are "Float-16 conversion instructions and the RDRND instruction". Float16 conversion would fit the F16C acronym in the error message.
The best option for us is then probably "-xHost -axCORE-AVX-I". This ensures (-x) that the code runs on the front end, but it will apply code optimizations (-ax) for the compute nodes.
-erik
users@lists.einsteintoolkit.org