#1023: Improve OpenMP parallelisation of SummationByParts -----------------------------------+---------------------------------------- Reporter: eschnett | Owner: Type: enhancement | Status: new Priority: major | Milestone: Component: EinsteinToolkit thorn | Version: Keywords: | -----------------------------------+---------------------------------------- The Intel compiler does not handle workshare constructs well. The attached patch replaces them by explicit loops, which execute faster. This makes a measurable difference on Hopper with 24 OpenMP threads.
#1023: Improve OpenMP parallelisation of SummationByParts ------------------------------------+--------------------------------------- Reporter: eschnett | Owner: Type: enhancement | Status: review Priority: major | Milestone: Component: EinsteinToolkit thorn | Version: Resolution: | Keywords: ------------------------------------+--------------------------------------- Changes (by eschnett):
* status: new => review
Old description:
The Intel compiler does not handle workshare constructs well. The attached patch replaces them by explicit loops, which execute faster. This makes a measurable difference on Hopper with 24 OpenMP threads.
New description:
The Intel compiler does not handle workshare constructs well. The attached patch replaces them by explicit loops, which execute faster. This makes a measurable difference on Hopper with 24 OpenMP threads.
This only modifies one operator; other operators could be treated in the same way.
--
#1023: Improve OpenMP parallelisation of SummationByParts ------------------------------------+--------------------------------------- Reporter: eschnett | Owner: Type: enhancement | Status: reviewed_ok Priority: major | Milestone: Component: EinsteinToolkit thorn | Version: Resolution: | Keywords: ------------------------------------+--------------------------------------- Changes (by knarf):
* status: review => reviewed_ok
Comment:
The patch looks ok. I didn't check all the indices really carefully (due to the length of the patch) and didn't run testsuites. Assuming tests show no difference between both versions using multiple threads I think it is ok to commit this. I'll leave testing to Erik. :)
#1023: Improve OpenMP parallelisation of SummationByParts ------------------------------------+--------------------------------------- Reporter: eschnett | Owner: Type: enhancement | Status: closed Priority: major | Milestone: Component: EinsteinToolkit thorn | Version: Resolution: fixed | Keywords: ------------------------------------+--------------------------------------- Changes (by eschnett):
* status: reviewed_ok => closed * resolution: => fixed
Comment:
Applied.
trac@lists.einsteintoolkit.org