hi all,
i've noticed that my runs (using latest ET release) with CarpetRegrid2 exhibit a significant increase in memory during runtime. this seems to happen immediately after some non-trivial regridding operation is done. the increase is steady, and at some point i run out of memory and the simulation crashes. this is happening both on my workstation (running Ubuntu 18.04) as well as our local cluster (running Debian 9). i was wondering if someone has seen something like this?
i have not seen this happen for simulations without CarpetRegrid2. i show below some relevant portions of the stdout file for a standard inspiral BH run (note the last column--maxrss_mb):
------------------------------------------------------------------------------------ Iteration Time | *me_per_hour | LEANBSSNMOL::conf_fac | *TISTICS::maxrss_mb | | minimum maximum | minimum maximum ------------------------------------------------------------------------------------ 0 0.000 | 0.0000000 | 0.2213400 0.9977828 | 1057 1359 4 0.025 | 1.8911444 | 0.2213388 0.9977828 | 1060 1361 8 0.050 | 2.8049414 | 0.2213330 0.9977828 | 1060 1361 12 0.075 | 3.2859229 | 0.2213195 0.9977828 | 1060 1361 16 0.100 | 3.6219375 | 0.2212959 0.9977828 | 1061 1361 20 0.125 | 3.7521230 | 0.2212596 0.9977828 | 1064 1361 24 0.150 | 3.9448186 | 0.2212081 0.9977828 | 1064 1361 28 0.175 | 4.0652624 | 0.2211358 0.9977828 | 1064 1361 32 0.200 | 4.1492619 | 0.2210378 0.9977828 | 1064 1361 36 0.225 | 4.1272708 | 0.2209119 0.9977828 | 1064 1361 40 0.250 | 4.2206476 | 0.2207447 0.9977828 | 1064 1361 44 0.275 | 4.2801401 | 0.2205343 0.9977828 | 1064 1361 48 0.300 | 4.3460198 | 0.2202766 0.9977828 | 1064 1361 52 0.325 | 4.3485616 | 0.2199651 0.9977828 | 1064 1361 56 0.350 | 4.4054911 | 0.2195938 0.9977828 | 1065 1361 60 0.375 | 4.4398502 | 0.2191577 0.9977828 | 1065 1361 64 0.400 | 4.3724216 | 0.2186524 0.9977828 | 1065 1361
(...)
INFO (QuasiLocalMeasures): Weinberg angular momentum z: 0.405200 384 2.400 | 4.2719649 | 0.3525426 0.9977828 | 1068 1379 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (Carpet): Grid structure (superregions, grid points): [0][0][0] exterior: [0,0,0] : [165,324,165] ([166,325,166] + PADDING) 8955700 [1][0][0] exterior: [3,155,3] : [179,493,175] ([177,339,173] + PADDING) 10380519 [2][0][0] exterior: [9,558,9] : [108,738,101] ([100,181,93] + PADDING) 1683300 [3][0][0] exterior: [21,1206,21] : [128,1386,113] ([108,181,93] + PADDING) 1817964 [4][0][0] exterior: [45,2521,45] : [147,2663,117] ([103,143,73] + PADDING) 1075217 [5][0][0] exterior: [93,5111,93] : [225,5257,165] ([133,147,73] + PADDING) 1427223 [6][0][0] exterior: [242,10307,189] : [380,10445,261] ([139,139,73] + PADDING) 1410433 [7][0][0] exterior: [554,20684,381] : [692,20822,453] ([139,139,73] + PADDING) 1410433 INFO (Carpet): Grid structure (superregions, coordinates): [0][0][0] exterior: [-4.800000000000001,-259.199999999999989,-4.800000000000001] : [259.199999999999989,259.199999999999989,259.199999999999989] : [1.600000000000000,1.600000000000000,1.600000000000000] [1][0][0] exterior: [-2.400000000000000,-135.199999999999989,-2.400000000000000] : [138.400000000000006,135.199999999999989,135.199999999999989] : [0.800000000000000,0.800000000000000,0.800000000000000] [2][0][0] exterior: [-1.200000000000001,-36.000000000000000,-1.200000000000001] : [38.400000000000006,36.000000000000000,35.600000000000009] : [0.400000000000000,0.400000000000000,0.400000000000000] [3][0][0] exterior: [-0.600000000000001,-18.000000000000000,-0.600000000000001] : [20.800000000000001,18.000000000000000,17.800000000000001] : [0.200000000000000,0.200000000000000,0.200000000000000] [4][0][0] exterior: [-0.300000000000001,-7.100000000000023,-0.300000000000001] : [9.900000000000000,7.099999999999966,6.900000000000000] : [0.100000000000000,0.100000000000000,0.100000000000000] [5][0][0] exterior: [-0.150000000000000,-3.650000000000006,-0.150000000000000] : [6.449999999999999,3.649999999999977,3.449999999999999] : [0.050000000000000,0.050000000000000,0.050000000000000] [6][0][0] exterior: [1.250000000000000,-1.525000000000034,-0.075000000000000] : [4.699999999999999,1.925000000000011,1.725000000000000] : [0.025000000000000,0.025000000000000,0.025000000000000] [7][0][0] exterior: [2.125000000000000,-0.650000000000034,-0.037500000000001] : [3.850000000000000,1.074999999999989,0.862500000000000] : [0.012500000000000,0.012500000000000,0.012500000000000] INFO (Carpet): Global grid structure statistics: INFO (Carpet): GF: rhs: 2551k active, 3661k owned (+44%), 6190k total (+69%), 320 steps/time INFO (Carpet): GF: vars: 161, pts: 3631M active, 4271M owned (+18%), 6261M total (+47%), 1.0 comp/proc INFO (Carpet): GA: vars: 1044, pts: 82M active, 82M total (+0%) INFO (Carpet): Total required memory: 50.821 GByte (for GAs and currently active GFs) INFO (Carpet): Load balance: min avg max sdv max/avg-1 INFO (Carpet): Level 0: 25M 27M 31M 1M owned 12% INFO (Carpet): Level 1: 31M 34M 36M 1M owned 6% INFO (Carpet): Level 2: 5M 5M 6M 0M owned 8% INFO (Carpet): Level 3: 5M 6M 6M 0M owned 8% INFO (Carpet): Level 4: 3M 3M 4M 0M owned 8% INFO (Carpet): Level 5: 4M 4M 5M 0M owned 9% INFO (Carpet): Level 6: 4M 5M 5M 0M owned 4% INFO (Carpet): Level 7: 4M 5M 5M 0M owned 4% 388 2.425 | 4.1930426 | 0.3508349 0.9977828 | 1149 1450 392 2.450 | 4.2022566 | 0.3491189 0.9977828 | 1149 1450 396 2.475 | 4.2091891 | 0.3471720 0.9977828 | 1149 1450 ------------------------------------------------------------------------------------ Iteration Time | *me_per_hour | LEANBSSNMOL::conf_fac | *TISTICS::maxrss_mb | | minimum maximum | minimum maximum ------------------------------------------------------------------------------------ 400 2.500 | 4.2177082 | 0.3450791 0.9977828 | 1149 1450 404 2.525 | 4.2195973 | 0.3429667 0.9977828 | 1149 1450
(...)
1344 8.400 | 4.4943523 | 0.3648129 0.9977828 | 1156 1455 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 0 INFO (CarpetRegrid2): Enforcing grid structure properties, iteration 1 INFO (Carpet): Grid structure (superregions, grid points): [0][0][0] exterior: [0,0,0] : [165,324,165] ([166,325,166] + PADDING) 8955700 [1][0][0] exterior: [3,154,3] : [178,494,175] ([176,341,173] + PADDING) 10382768 [2][0][0] exterior: [9,556,9] : [108,740,101] ([100,185,93] + PADDING) 1720500 [3][0][0] exterior: [21,1202,21] : [127,1390,113] ([107,189,93] + PADDING) 1880739 [4][0][0] exterior: [45,2512,45] : [145,2672,117] ([101,161,73] + PADDING) 1187053 [5][0][0] exterior: [93,5094,93] : [110,5135,165] ([18,42,73] + PADDING) 55188 [5][0][1] exterior: [93,5136,93] : [220,5274,165] ([128,139,73] + PADDING) 1298816 [6][0][0] exterior: [234,10342,189] : [372,10480,261] ([139,139,73] + PADDING) 1410433 [7][0][0] exterior: [537,20752,381] : [675,20890,453] ([139,139,73] + PADDING) 1410433 INFO (Carpet): Grid structure (superregions, coordinates): [0][0][0] exterior: [-4.800000000000001,-259.199999999999989,-4.800000000000001] : [259.199999999999989,259.199999999999989,259.199999999999989] : [1.600000000000000,1.600000000000000,1.600000000000000] [1][0][0] exterior: [-2.400000000000000,-136.000000000000000,-2.400000000000000] : [137.599999999999994,136.000000000000000,135.199999999999989] : [0.800000000000000,0.800000000000000,0.800000000000000] [2][0][0] exterior: [-1.200000000000001,-36.800000000000011,-1.200000000000001] : [38.400000000000006,36.800000000000011,35.600000000000009] : [0.400000000000000,0.400000000000000,0.400000000000000] [3][0][0] exterior: [-0.600000000000001,-18.800000000000011,-0.600000000000001] : [20.600000000000001,18.800000000000011,17.800000000000001] : [0.200000000000000,0.200000000000000,0.200000000000000] [4][0][0] exterior: [-0.300000000000001,-8.000000000000000,-0.300000000000001] : [9.699999999999999,8.000000000000000,6.900000000000000] : [0.100000000000000,0.100000000000000,0.100000000000000] [5][0][0] exterior: [-0.150000000000000,-4.500000000000000,-0.150000000000000] : [0.699999999999999,-2.449999999999989,3.449999999999999] : [0.050000000000000,0.050000000000000,0.050000000000000] [5][0][1] exterior: [-0.150000000000000,-2.400000000000034,-0.150000000000000] : [6.199999999999999,4.500000000000000,3.449999999999999] : [0.050000000000000,0.050000000000000,0.050000000000000] [6][0][0] exterior: [1.050000000000000,-0.650000000000034,-0.075000000000000] : [4.500000000000000,2.800000000000011,1.725000000000000] : [0.025000000000000,0.025000000000000,0.025000000000000] [7][0][0] exterior: [1.912500000000000,0.199999999999989,-0.037500000000001] : [3.637499999999999,1.925000000000011,0.862500000000000] : [0.012500000000000,0.012500000000000,0.012500000000000] INFO (Carpet): Global grid structure statistics: INFO (Carpet): GF: rhs: 2544k active, 3659k owned (+44%), 6187k total (+69%), 320 steps/time INFO (Carpet): GF: vars: 161, pts: 3644M active, 4290M owned (+18%), 6288M total (+47%), 1.0 comp/proc INFO (Carpet): GA: vars: 1044, pts: 82M active, 82M total (+0%) INFO (Carpet): Total required memory: 51.040 GByte (for GAs and currently active GFs) INFO (Carpet): Load balance: min avg max sdv max/avg-1 INFO (Carpet): Level 0: 25M 27M 31M 1M owned 12% INFO (Carpet): Level 1: 31M 34M 35M 1M owned 4% INFO (Carpet): Level 2: 5M 5M 6M 0M owned 6% INFO (Carpet): Level 3: 5M 6M 6M 0M owned 8% INFO (Carpet): Level 4: 3M 4M 4M 0M owned 10% INFO (Carpet): Level 5: 3M 4M 5M 0M owned 11% INFO (Carpet): Level 6: 4M 5M 5M 0M owned 4% INFO (Carpet): Level 7: 4M 5M 5M 0M owned 4% 1348 8.425 | 4.4749366 | 0.3620783 0.9977828 | 1417 1726 1352 8.450 | 4.4793418 | 0.3593239 0.9977828 | 1417 1727 1356 8.475 | 4.4831132 | 0.3565493 0.9977828 | 1417 1727 ------------------------------------------------------------------------------------ Iteration Time | *me_per_hour | LEANBSSNMOL::conf_fac | *TISTICS::maxrss_mb | | minimum maximum | minimum maximum ------------------------------------------------------------------------------------ 1360 8.500 | 4.4873776 | 0.3537546 0.9977828 | 1417 1727 1364 8.525 | 4.4896835 | 0.3509394 0.9977828 | 1417 1727
and the pattern continues such that at iteration 9000 we're at maximum(maxrss_mb) = 2722... is this normal, or expected? it becomes very inconvenient since, for some high-resolutions runs, i inevitably run out of memory.
thanks, Miguel
On 19 Jul 2018, at 11:58, Miguel Zilhão miguel.zilhao.nogueira@tecnico.ulisboa.pt wrote:
hi all,
i've noticed that my runs (using latest ET release) with CarpetRegrid2 exhibit a significant increase in memory during runtime. this seems to happen immediately after some non-trivial regridding operation is done. the increase is steady, and at some point i run out of memory and the simulation crashes. this is happening both on my workstation (running Ubuntu 18.04) as well as our local cluster (running Debian 9). i was wondering if someone has seen something like this?
i have not seen this happen for simulations without CarpetRegrid2. i show below some relevant portions of the stdout file for a standard inspiral BH run (note the last column--maxrss_mb):
This could be caused by memory fragmentation due to all the freeing and mallocing that happens during regridding when the sizes of the grids change. Can you try using tcmalloc or jemalloc instead of glibc malloc and reporting back? One workaround could be to run shorter simulations (i.e. set a walltime of 12 h instead of 24 h).
hi Ian,
i've noticed that my runs (using latest ET release) with CarpetRegrid2 exhibit a significant increase in memory during runtime. this seems to happen immediately after some non-trivial regridding operation is done. the increase is steady, and at some point i run out of memory and the simulation crashes. this is happening both on my workstation (running Ubuntu 18.04) as well as our local cluster (running Debian 9). i was wondering if someone has seen something like this?
i have not seen this happen for simulations without CarpetRegrid2. i show below some relevant portions of the stdout file for a standard inspiral BH run (note the last column--maxrss_mb):
This could be caused by memory fragmentation due to all the freeing and mallocing that happens during regridding when the sizes of the grids change. Can you try using tcmalloc or jemalloc instead of glibc malloc and reporting back? One workaround could be to run shorter simulations (i.e. set a walltime of 12 h instead of 24 h).
thanks for your reply. in one of my cases, for the resolution used and the available memory, i was out of memory quite quickly -- within 6 hours or so... so unfortunately it becomes a bit impractical for large simulations...
what would i need to do in order to use tcmalloc or jemalloc?
thanks, Miguel
On 19 Jul 2018, at 17:14, Miguel Zilhão miguel.zilhao.nogueira@tecnico.ulisboa.pt wrote:
hi Ian,
i've noticed that my runs (using latest ET release) with CarpetRegrid2 exhibit a significant increase in memory during runtime. this seems to happen immediately after some non-trivial regridding operation is done. the increase is steady, and at some point i run out of memory and the simulation crashes. this is happening both on my workstation (running Ubuntu 18.04) as well as our local cluster (running Debian 9). i was wondering if someone has seen something like this?
i have not seen this happen for simulations without CarpetRegrid2. i show below some relevant portions of the stdout file for a standard inspiral BH run (note the last column--maxrss_mb):
This could be caused by memory fragmentation due to all the freeing and mallocing that happens during regridding when the sizes of the grids change. Can you try using tcmalloc or jemalloc instead of glibc malloc and reporting back? One workaround could be to run shorter simulations (i.e. set a walltime of 12 h instead of 24 h).
thanks for your reply. in one of my cases, for the resolution used and the available memory, i was out of memory quite quickly -- within 6 hours or so... so unfortunately it becomes a bit impractical for large simulations...
what would i need to do in order to use tcmalloc or jemalloc?
I have used tcmalloc. I think you will need the following:
- Install tcmalloc (https://github.com/gperftools/gperftools), and libunwind, which it depends on. - In your optionlist, link with tcmalloc. I have
LDFLAGS = -rdynamic -L/home/ianhin/software/gperftools-2.1/lib -Wl,-rpath,/home/ianhin/software/gperftools-2.1/lib -ltcmalloc
This should be sufficient I think for tcmalloc to be used instead of glibc malloc. Try this out, and see if things are better. I also have a thorn which hooks into the tcmalloc API. You can get it from
https://bitbucket.org/ianhinder/tcmalloc
It's very much a work in progress, and probably has some hard-coded assumptions in it. You can set Cactus parameters to:
1. Report memory statistics periodically 2. Release memory back to the OS periodically 3. Output a memory profile periodically
Let us know how it goes!
hi Ian and all,
This could be caused by memory fragmentation due to all the freeing and mallocing that happens during regridding when the sizes of the grids change. Can you try using tcmalloc or jemalloc instead of glibc malloc and reporting back? One workaround could be to run shorter simulations (i.e. set a walltime of 12 h instead of 24 h).
thanks for your reply. in one of my cases, for the resolution used and the available memory, i was out of memory quite quickly -- within 6 hours or so... so unfortunately it becomes a bit impractical for large simulations...
what would i need to do in order to use tcmalloc or jemalloc?
I have used tcmalloc. I think you will need the following:
- Install tcmalloc (https://github.com/gperftools/gperftools), and libunwind, which it depends on.
- In your optionlist, link with tcmalloc. I have
LDFLAGS = -rdynamic -L/home/ianhin/software/gperftools-2.1/lib -Wl,-rpath,/home/ianhin/software/gperftools-2.1/lib -ltcmalloc
This should be sufficient I think for tcmalloc to be used instead of glibc malloc. Try this out, and see if things are better. I also have a thorn which hooks into the tcmalloc API. You can get it from
thanks a lot for these pointers. i've tried it out, though i've used tcmalloc from ubuntu's repositories and therefore compiled ET with -ltcmalloc_minimal. i don't know whether this makes a difference, but from the trial run that i'm doing i so far seem to see the same memory increase i had seen before...
is there anything else that can be tried to try to pinpoint this issue? it seems to be serious... i looked for an open ticket but didn't find anything. shall i submit one?
thanks, Miguel
On 23 Jul 2018, at 15:00, Miguel Zilhão miguel.zilhao.nogueira@tecnico.ulisboa.pt wrote:
hi Ian and all,
This could be caused by memory fragmentation due to all the freeing and mallocing that happens during regridding when the sizes of the grids change. Can you try using tcmalloc or jemalloc instead of glibc malloc and reporting back? One workaround could be to run shorter simulations (i.e. set a walltime of 12 h instead of 24 h).
thanks for your reply. in one of my cases, for the resolution used and the available memory, i was out of memory quite quickly -- within 6 hours or so... so unfortunately it becomes a bit impractical for large simulations...
what would i need to do in order to use tcmalloc or jemalloc?
I have used tcmalloc. I think you will need the following:
- Install tcmalloc (https://github.com/gperftools/gperftools), and libunwind, which it depends on.
- In your optionlist, link with tcmalloc. I have
LDFLAGS = -rdynamic -L/home/ianhin/software/gperftools-2.1/lib -Wl,-rpath,/home/ianhin/software/gperftools-2.1/lib -ltcmalloc This should be sufficient I think for tcmalloc to be used instead of glibc malloc. Try this out, and see if things are better. I also have a thorn which hooks into the tcmalloc API. You can get it from
thanks a lot for these pointers. i've tried it out, though i've used tcmalloc from ubuntu's repositories and therefore compiled ET with -ltcmalloc_minimal. i don't know whether this makes a difference, but from the trial run that i'm doing i so far seem to see the same memory increase i had seen before...
Hi,
Did you use the tcmalloc thorn and the parameters to make it release memory back to the OS after each regridding?
is there anything else that can be tried to try to pinpoint this issue? it seems to be serious... i looked for an open ticket but didn't find anything. shall i submit one?
You can look at the Carpet::memory_procs variable (in Carpet/Carpet/interface.ccl). This will tell you how much memory Carpet has allocated for different things. If one of these is growing during the run, but resets to a lower value after checkpoint recovery, then that suggests a memory leak.
Hmm. Now this is coming back to me. I just searched my email, and I found a draft email that I never sent from 2015 with subject "Memory leak in Carpet?". Here it is:
Does anyone have any reason to suspect that something in Carpet might be leaking memory? I have a fairly straightforward QC0 simulation, based on qc0-mclachlan, and it looks like something is leaking memory.
I have looked at several diagnostics.
– carpet::grid_functions: this contains the amount of memory Carpet thinks is allocated for grid functions. Since I have regridding, this is not in general going to be constant, but it turns out that it remains approximate constant at about 5 GB per process (average and maximum across processes are about the same). Each process has 12 GB available on Datura. So I should be well within the memory limits of the machine, by more than a factor of 2.
– The run crashes with signal 9 during regridding at some point; this is probably the OOM killer.
– SystemStatistics::swap_used_mb starts to grow after the first regridding, and seems to grow linearly throughout the run. The crash time corresponds to it hitting about 18 GB, which is the maximum swap configured on the node.
– SystemStatistics::arena: This is the 'arena' field from mallinfo (http://man7.org/linux/man-pages/man3/mallinfo.3.html http://man7.org/linux/man-pages/man3/mallinfo.3.html) which is supposed to give 'Non-mmapped space allocated (bytes)'. This suffers from being stored in a 32 bit integer in the mallinfo structure (https://sourceware.org/ml/libc-alpha/2014-11/msg00431.html https://sourceware.org/ml/libc-alpha/2014-11/msg00431.html), but I have adjusted it manually by adding 4 GB when it looks like it is dropping unphysically. This shows that the amount of non-mmapped memory allocated is increasing on the order of 1 GB on each regridding, and not between regriddings. Firstly, the amount of memory allocated shouldn't be increasing so much, since carpet::grid_functions remains approximately constant. Secondly, I thought that we were supposed to be using mmap for grid function data, so why do we have such a large amount of non-mmapped memory?
– SystemStatistics::hblkhd: This is the hblkhd field from mallinfo, which is supposed to be "Space allocated in mmapped regions (bytes)". This increases to about 2 GB after a couple of regriddings, but then stays roughly constant at 2 GB, which seems fine, apart from the fact that I had hoped that mmap was being used for all the gridfunction data
I suspect I got distracted doing my own investigations while I was writing it, and then lost track of it. Typically, I think I run with excess memory for each run, so that by 24 h, it hasn't grown too badly.
Further searching finds an email from someone else saying that they had memory leak problems and asking me about it. My reply was:
On 5 May 2015, at 17:14, Ian Hinder ian.hinder@aei.mpg.de wrote:
Hi,
I have added Erik and Roland to CC, as we have been discussing this; I hope this is OK.
It sounds very similar. I am running on several hundred cores (<600) and the simulations often fail with OOM errors after less than a day. First the RSS grows, then the swap, then the OOM-killer kills it. I have observed this both on Datura and Hydra. Stopping the simulations and recovering from a checkpoint usually fixes the problem, and it runs on for another half day or so. I have done a fair amount of work on this, so I will summarise here.
Monitor process RSS
The process resident set size is the amount of address space which is currently mapped into physical memory by the OS. Thorn SystemStatistics can be used to measure this and put it into a Cactus variable, which can then be reduced across processes. I use:
IOBasic::outInfo_every = 1 IOBasic::outInfo_reductions = "maximum" IOBasic::outInfo_vars = " SystemStatistics::maxrss_mb SystemStatistics::swap_used_mb Carpet::gridfunctions "
and
IOScalar::outScalar_every = 128 IOScalar::outScalar_vars = " SystemStatistics::process_memory_mb Carpet::memory_procs "
SystemStatistics calls its variable "maxrss" but it should actually be called "rss", as that is what is output. maxrss is also available from the OS, and would give the maximum the RSS had ever been during the process lifetime.
Carpet::gridfunctions (in Carpet::memory_procs) measures the amount of memory Carpet has allocated in gridfunctions. For me, this remains essentially flat, whereas maxrss grows after each regridding until it reaches the maximum available, then the swap starts to grow. This indicates that the problem is not due to Carpet allocating more and more grid points due to grids changing size. It could be due to failing to free allocated memory (a leak) or freed data taking up space which cannot be used for further allocations or returned to the OS (fragmentation).
Terminate and checkpoint on OOM
I have a local change to SystemStatistics which adds parameters for maximum values of RSS and swap usage, above which it calls CCTK_TerminateNext, so if you have checkpoint_on_terminate, you get a clean termination and can continue the run without losing too much CPU time. I have been running with this for a couple of weeks now, and it works as advertised. I have a branch with this on, but I just realised it conflicts with a change Erik made. If you want this, let me know and I will sort it out.
Memory profiling
The malloc implementation in glibc provides no usable statistics. mallinfo is limited to 32 bit integers, which overflow for 64 bit systems. malloc_info, at least in the version on datura, doesn't include memory allocated via mmap. Useless. Instead, you need to use an external memory profiler to see what is going on. I have used "igprof" successfully, and this shows me that there is no "leak" of allocated memory corresponding to the increase in RSS. i.e. the problem is not caused by forgetting to free something. This suggests that the problem is fragmentation, where malloc has unallocated blocks of memory which it does not or cannot return to the OS. Malloc allocates memory in two ways: either in its main heap, or by allocating anonymous mmap regions. I had thought that only the latter could be fully returned to the OS, but this is not true. Any region of address space can be marked as unused (internally via the madvise(MADV_DONTNEED) system call) and a malloc implementation can do this on regions of its address space which have been freed. If such regions are too small (smaller than a page), then they could accumulate and not be returned to the OS.
Alternative malloc implementations
At the suggestion of Roland, I tried using the tcmalloc (http://gperftools.googlecode.com/git/doc/tcmalloc.html http://gperftools.googlecode.com/git/doc/tcmalloc.html) library, which is a drop-in replacement for glibc malloc which is part of gperftools. This works fairly easily. You can compile the "minimal" version with no dependencies and then modify your optionlist:
CPPFLAGS = -I/home/rhaas/software/gperftools-2.1/include/gperftools LDFLAGS = -L/home/rhaas/software/gperftools-2.1/lib -Wl,-rpath,/home/rhaas/software/gperftools-2.1/lib -ltcmalloc_minimal
I found in one example case that this reduced the RSS process growth, so I am now using it for all my simulations. However, I still run into the same problem eventually, so it might be that it makes it better but doesn't solve it completely.
tcmalloc has an introspection interface which is presumably not as useless as glibc's malloc: http://gperftools.googlecode.com/git/doc/tcmalloc.html http://gperftools.googlecode.com/git/doc/tcmalloc.html. I haven't tried this yet.
Checkpoint recovery
I noticed from the igprof profile that there are 11000 allocations (and frees) during checkpoint recovery on one process, all from the HDF5 uncompression routine. This is the "deflate" filter. When it decompresses a dataset, it allocates a buffer, initially sized the same as the compressed dataset (really dumb, as it will always need to be bigger). It then uncompresses into the buffer, "realloc"ing the buffer to twice the size each time it runs out of space. You can imagine that this might cause a lot of fragmentation. There is no tunable parameter, but we could modify the code (it's in https://svn.hdfgroup.uiuc.edu/hdf5/tags/hdf5-1_8_12/src/H5Zdeflate.c https://svn.hdfgroup.uiuc.edu/hdf5/tags/hdf5-1_8_12/src/H5Zdeflate.c) to use a much larger starting buffer size, in the hope that this reduces the number of reallocs, and hence the amount of fragmentation. This wouldn't help the accumulated RSS, but it would probably produce a one-off decrease in the amount of fragmentation. I am currently not using periodic checkpointing, so I don't know if the compression routine has the same problem. Probably not, since it knows the output buffer size has to be smaller than the input buffer size. Apparently Frank Löffler also modified this routine, which solved some of his problems of running out of memory during recovery. Another alternative would be to disable checkpoint compression.
To see if you are suffering from the same problem, I think the quickest way would be to link against tcmalloc and use MallocExtension::instance()->GetNumericProperty(property_name, value) from tcmalloc to read off the generic.current_allocated_bytes property (Number of bytes used by the application. This will not typically match the memory use reported by the OS, because it does not include TCMalloc overhead or memory fragmentation). You could also look at the other properties they provide. Then compare this with the process RSS from systemstatistics, and Carpet's gridfunctions variable, and check to see if you actually have a memory leak, or if you are suffering from fragmentation.
This brings to mind another question: are you using HDF5 compression? If so, do you see the same problem if you switch it off?
And finally: do you get this growth on the very first segment of a simulation, or only on subsequent segments? I am thinking that checkpoint *recovery* severely fragments the heap, especially with compression, and this somehow causes growth in the RSS with each regridding.
So, while I had forgotten about all this, it turns out that I had actually thought quite a lot about it :)
hi Ian,
thanks again for your thorough reply! i've checked the low resolution run that i have (and which ran successfully) and the pattern that i observe is that maxrss typically grows after each regridding. the increase in maxrss is not monotonic, though, as it does go down at times. then, after the BHs merge and the regridding stops, maxrss settles down (to a value still a bit higher than the maxrss at t=0). this is all for the first segment of the simulation, as i haven't done any checkpointing for this run.
so i guess i don't have enough data to conclude whether what i'm observing is indeed a memory leak, or if it's just Carpet's regridding algorithm doing what it's supposed to do. in any case, i guess what surprises me is that, at times, the memory consumption can be bigger than what it was at the beginning of the run by a factor of 1.6, which can easily lead to an out-of-memory situation...
unfortunately these days i don't have the time to investigate this any further...
many thanks, Miguel
On 23/07/2018 16:06, ian.hinder@aei.mpg.de wrote:
On 23 Jul 2018, at 15:00, Miguel Zilhão <miguel.zilhao.nogueira@tecnico.ulisboa.pt mailto:miguel.zilhao.nogueira@tecnico.ulisboa.pt> wrote:
hi Ian and all,
This could be caused by memory fragmentation due to all the freeing and mallocing that happens during regridding when the sizes of the grids change. Can you try using tcmalloc or jemalloc instead of glibc malloc and reporting back? One workaround could be to run shorter simulations (i.e. set a walltime of 12 h instead of 24 h).
thanks for your reply. in one of my cases, for the resolution used and the available memory, i was out of memory quite quickly -- within 6 hours or so... so unfortunately it becomes a bit impractical for large simulations...
what would i need to do in order to use tcmalloc or jemalloc?
I have used tcmalloc. I think you will need the following:
- Install tcmalloc (https://github.com/gperftools/gperftools), and libunwind, which it depends on.
- In your optionlist, link with tcmalloc. I have
LDFLAGS = -rdynamic -L/home/ianhin/software/gperftools-2.1/lib -Wl,-rpath,/home/ianhin/software/gperftools-2.1/lib -ltcmalloc This should be sufficient I think for tcmalloc to be used instead of glibc malloc. Try this out, and see if things are better. I also have a thorn which hooks into the tcmalloc API. You can get it from
thanks a lot for these pointers. i've tried it out, though i've used tcmalloc from ubuntu's repositories and therefore compiled ET with -ltcmalloc_minimal. i don't know whether this makes a difference, but from the trial run that i'm doing i so far seem to see the same memory increase i had seen before...
Hi,
Did you use the tcmalloc thorn and the parameters to make it release memory back to the OS after each regridding?
is there anything else that can be tried to try to pinpoint this issue? it seems to be serious... i looked for an open ticket but didn't find anything. shall i submit one?
You can look at the Carpet::memory_procs variable (in Carpet/Carpet/interface.ccl). This will tell you how much memory Carpet has allocated for different things. If one of these is growing during the run, but resets to a lower value after checkpoint recovery, then that suggests a memory leak.
Hmm. Now this is coming back to me. I just searched my email, and I found a draft email that I never sent from 2015 with subject "Memory leak in Carpet?". Here it is:
Does anyone have any reason to suspect that something in Carpet might be leaking memory? I have a fairly straightforward QC0 simulation, based on qc0-mclachlan, and it looks like something is leaking memory.
I have looked at several diagnostics.
– carpet::grid_functions: this contains the amount of memory Carpet thinks is allocated for grid functions. Since I have regridding, this is not in general going to be constant, but it turns out that it remains approximate constant at about 5 GB per process (average and maximum across processes are about the same). Each process has 12 GB available on Datura. So I should be well within the memory limits of the machine, by more than a factor of 2.
– The run crashes with signal 9 during regridding at some point; this is probably the OOM killer. – SystemStatistics::swap_used_mb starts to grow after the first regridding, and seems to grow linearly throughout the run. The crash time corresponds to it hitting about 18 GB, which is the maximum swap configured on the node.
– SystemStatistics::arena: This is the 'arena' field from mallinfo (http://man7.org/linux/man-pages/man3/mallinfo.3.html) which is supposed to give 'Non-mmapped space allocated (bytes)'. This suffers from being stored in a 32 bit integer in the mallinfo structure (https://sourceware.org/ml/libc-alpha/2014-11/msg00431.html), but I have adjusted it manually by adding 4 GB when it looks like it is dropping unphysically. This shows that the amount of non-mmapped memory allocated is increasing on the order of 1 GB on each regridding, and not between regriddings. Firstly, the amount of memory allocated shouldn't be increasing so much, since carpet::grid_functions remains approximately constant. Secondly, I thought that we were supposed to be using mmap for grid function data, so why do we have such a large amount of non-mmapped memory?
– SystemStatistics::hblkhd: This is the hblkhd field from mallinfo, which is supposed to be "Space allocated in mmapped regions (bytes)". This increases to about 2 GB after a couple of regriddings, but then stays roughly constant at 2 GB, which seems fine, apart from the fact that I had hoped that mmap was being used for all the gridfunction data
I suspect I got distracted doing my own investigations while I was writing it, and then lost track of it. Typically, I think I run with excess memory for each run, so that by 24 h, it hasn't grown too badly.
Further searching finds an email from someone else saying that they had memory leak problems and asking me about it. My reply was:
On 5 May 2015, at 17:14, Ian Hinder <ian.hinder@aei.mpg.de mailto:ian.hinder@aei.mpg.de> wrote:
Hi,
I have added Erik and Roland to CC, as we have been discussing this; I hope this is OK.
It sounds very similar. I am running on several hundred cores (<600) and the simulations often fail with OOM errors after less than a day. First the RSS grows, then the swap, then the OOM-killer kills it. I have observed this both on Datura and Hydra. Stopping the simulations and recovering from a checkpoint usually fixes the problem, and it runs on for another half day or so. I have done a fair amount of work on this, so I will summarise here.
*Monitor process RSS*
The process resident set size is the amount of address space which is currently mapped into physical memory by the OS. Thorn SystemStatistics can be used to measure this and put it into a Cactus variable, which can then be reduced across processes. I use:
IOBasic::outInfo_every = 1 IOBasic::outInfo_reductions = "maximum" IOBasic::outInfo_vars = " SystemStatistics::maxrss_mb SystemStatistics::swap_used_mb Carpet::gridfunctions "
and
IOScalar::outScalar_every = 128 IOScalar::outScalar_vars = " SystemStatistics::process_memory_mb Carpet::memory_procs "
SystemStatistics calls its variable "maxrss" but it should actually be called "rss", as that is what is output. maxrss is also available from the OS, and would give the maximum the RSS had ever been during the process lifetime.
Carpet::gridfunctions (in Carpet::memory_procs) measures the amount of memory Carpet has allocated in gridfunctions. For me, this remains essentially flat, whereas maxrss grows after each regridding until it reaches the maximum available, then the swap starts to grow. This indicates that the problem is not due to Carpet allocating more and more grid points due to grids changing size. It could be due to failing to free allocated memory (a leak) or freed data taking up space which cannot be used for further allocations or returned to the OS (fragmentation).
*Terminate and checkpoint on OOM*
I have a local change to SystemStatistics which adds parameters for maximum values of RSS and swap usage, above which it calls CCTK_TerminateNext, so if you have checkpoint_on_terminate, you get a clean termination and can continue the run without losing too much CPU time. I have been running with this for a couple of weeks now, and it works as advertised. I have a branch with this on, but I just realised it conflicts with a change Erik made. If you want this, let me know and I will sort it out.
*Memory profiling*
The malloc implementation in glibc provides no usable statistics. mallinfo is limited to 32 bit integers, which overflow for 64 bit systems. malloc_info, at least in the version on datura, doesn't include memory allocated via mmap. Useless. Instead, you need to use an external memory profiler to see what is going on. I have used "igprof" successfully, and this shows me that there is no "leak" of allocated memory corresponding to the increase in RSS. i.e. the problem is not caused by forgetting to free something. This suggests that the problem is fragmentation, where malloc has unallocated blocks of memory which it does not or cannot return to the OS. Malloc allocates memory in two ways: either in its main heap, or by allocating anonymous mmap regions. I had thought that only the latter could be fully returned to the OS, but this is not true. Any region of address space can be marked as unused (internally via the madvise(MADV_DONTNEED) system call) and a malloc implementation can do this on regions of its address space which have been freed. If such regions are too small (smaller than a page), then they could accumulate and not be returned to the OS.
*Alternative malloc implementations*
At the suggestion of Roland, I tried using the tcmalloc (http://gperftools.googlecode.com/git/doc/tcmalloc.html) library, which is a drop-in replacement for glibc malloc which is part of gperftools. This works fairly easily. You can compile the "minimal" version with no dependencies and then modify your optionlist:
CPPFLAGS = -I/home/rhaas/software/gperftools-2.1/include/gperftools LDFLAGS = -L/home/rhaas/software/gperftools-2.1/lib -Wl,-rpath,/home/rhaas/software/gperftools-2.1/lib -ltcmalloc_minimal
I found in one example case that this reduced the RSS process growth, so I am now using it for all my simulations. However, I still run into the same problem eventually, so it might be that it makes it better but doesn't solve it completely.
tcmalloc has an introspection interface which is presumably not as useless as glibc's malloc: http://gperftools.googlecode.com/git/doc/tcmalloc.html. I haven't tried this yet.
*Checkpoint recovery*
I noticed from the igprof profile that there are 11000 allocations (and frees) during checkpoint recovery on one process, all from the HDF5 uncompression routine. This is the "deflate" filter. When it decompresses a dataset, it allocates a buffer, initially sized the same as the compressed dataset (really dumb, as it will always need to be bigger). It then uncompresses into the buffer, "realloc"ing the buffer to twice the size each time it runs out of space. You can imagine that this might cause a lot of fragmentation. There is no tunable parameter, but we could modify the code (it's in https://svn.hdfgroup.uiuc.edu/hdf5/tags/hdf5-1_8_12/src/H5Zdeflate.c) to use a much larger starting buffer size, in the hope that this reduces the number of reallocs, and hence the amount of fragmentation. This wouldn't help the accumulated RSS, but it would probably produce a one-off decrease in the amount of fragmentation. I am currently not using periodic checkpointing, so I don't know if the compression routine has the same problem. Probably not, since it knows the output buffer size has to be smaller than the input buffer size. Apparently Frank Löffler also modified this routine, which solved some of his problems of running out of memory during recovery. Another alternative would be to disable checkpoint compression.
To see if you are suffering from the same problem, I think the quickest way would be to link against tcmalloc and use MallocExtension::instance()->GetNumericProperty(property_name, value) from tcmalloc to read off the generic.current_allocated_bytes property (Number of bytes used by the application. This will not typically match the memory use reported by the OS, because it does not include TCMalloc overhead or memory fragmentation). You could also look at the other properties they provide. Then compare this with the process RSS from systemstatistics, and Carpet's gridfunctions variable, and check to see if you actually have a memory leak, or if you are suffering from fragmentation.
This brings to mind another question: are you using HDF5 compression? If so, do you see the same problem if you switch it off?
And finally: do you get this growth on the very first segment of a simulation, or only on subsequent segments? I am thinking that checkpoint *recovery* severely fragments the heap, especially with compression, and this somehow causes growth in the RSS with each regridding.
So, while I had forgotten about all this, it turns out that I had actually thought quite a lot about it :)
-- Ian Hinder https://ianhinder.net
hi all,
some more information regarding this. i've ran a simulation based on the provided parfile qc0-mclachlan.par. i'm attaching a crude plot with the memory consumption, as reported by systemstatistics-process_memory and carpet-memory_procs, as function of the iteration.
the jump at iteration ~2048 corresponds to the time where some non-trivial regridding occurred. if i'm interpreting the plot correctly, while the total memory consumption of the system (as reported by systemstatistics) increases, the memory that carpet reports to be using is actually decreasing. is this a sign that something is probably leaking memory?
in case it's useful, i'm also attaching the exact parameter file i've used.
thanks, Miguel
On 29/07/2018 21:12, Miguel Zilhão wrote:
hi Ian,
thanks again for your thorough reply! i've checked the low resolution run that i have (and which ran successfully) and the pattern that i observe is that maxrss typically grows after each regridding. the increase in maxrss is not monotonic, though, as it does go down at times. then, after the BHs merge and the regridding stops, maxrss settles down (to a value still a bit higher than the maxrss at t=0). this is all for the first segment of the simulation, as i haven't done any checkpointing for this run.
so i guess i don't have enough data to conclude whether what i'm observing is indeed a memory leak, or if it's just Carpet's regridding algorithm doing what it's supposed to do. in any case, i guess what surprises me is that, at times, the memory consumption can be bigger than what it was at the beginning of the run by a factor of 1.6, which can easily lead to an out-of-memory situation...
unfortunately these days i don't have the time to investigate this any further...
many thanks, Miguel
On 23/07/2018 16:06, ian.hinder@aei.mpg.de wrote:
On 23 Jul 2018, at 15:00, Miguel Zilhão <miguel.zilhao.nogueira@tecnico.ulisboa.pt mailto:miguel.zilhao.nogueira@tecnico.ulisboa.pt> wrote:
hi Ian and all,
This could be caused by memory fragmentation due to all the freeing and mallocing that happens during regridding when the sizes of the grids change. Can you try using tcmalloc or jemalloc instead of glibc malloc and reporting back? One workaround could be to run shorter simulations (i.e. set a walltime of 12 h instead of 24 h).
thanks for your reply. in one of my cases, for the resolution used and the available memory, i was out of memory quite quickly -- within 6 hours or so... so unfortunately it becomes a bit impractical for large simulations...
what would i need to do in order to use tcmalloc or jemalloc?
I have used tcmalloc. I think you will need the following:
- Install tcmalloc (https://github.com/gperftools/gperftools), and libunwind, which it depends on.
- In your optionlist, link with tcmalloc. I have
LDFLAGS = -rdynamic -L/home/ianhin/software/gperftools-2.1/lib -Wl,-rpath,/home/ianhin/software/gperftools-2.1/lib -ltcmalloc This should be sufficient I think for tcmalloc to be used instead of glibc malloc. Try this out, and see if things are better. I also have a thorn which hooks into the tcmalloc API. You can get it from
thanks a lot for these pointers. i've tried it out, though i've used tcmalloc from ubuntu's repositories and therefore compiled ET with -ltcmalloc_minimal. i don't know whether this makes a difference, but from the trial run that i'm doing i so far seem to see the same memory increase i had seen before...
Hi,
Did you use the tcmalloc thorn and the parameters to make it release memory back to the OS after each regridding?
is there anything else that can be tried to try to pinpoint this issue? it seems to be serious... i looked for an open ticket but didn't find anything. shall i submit one?
You can look at the Carpet::memory_procs variable (in Carpet/Carpet/interface.ccl). This will tell you how much memory Carpet has allocated for different things. If one of these is growing during the run, but resets to a lower value after checkpoint recovery, then that suggests a memory leak.
Hmm. Now this is coming back to me. I just searched my email, and I found a draft email that I never sent from 2015 with subject "Memory leak in Carpet?". Here it is:
Does anyone have any reason to suspect that something in Carpet might be leaking memory? I have a fairly straightforward QC0 simulation, based on qc0-mclachlan, and it looks like something is leaking memory.
I have looked at several diagnostics.
– carpet::grid_functions: this contains the amount of memory Carpet thinks is allocated for grid functions. Since I have regridding, this is not in general going to be constant, but it turns out that it remains approximate constant at about 5 GB per process (average and maximum across processes are about the same). Each process has 12 GB available on Datura. So I should be well within the memory limits of the machine, by more than a factor of 2.
– The run crashes with signal 9 during regridding at some point; this is probably the OOM killer. – SystemStatistics::swap_used_mb starts to grow after the first regridding, and seems to grow linearly throughout the run. The crash time corresponds to it hitting about 18 GB, which is the maximum swap configured on the node.
– SystemStatistics::arena: This is the 'arena' field from mallinfo (http://man7.org/linux/man-pages/man3/mallinfo.3.html) which is supposed to give 'Non-mmapped space allocated (bytes)'. This suffers from being stored in a 32 bit integer in the mallinfo structure (https://sourceware.org/ml/libc-alpha/2014-11/msg00431.html), but I have adjusted it manually by adding 4 GB when it looks like it is dropping unphysically. This shows that the amount of non-mmapped memory allocated is increasing on the order of 1 GB on each regridding, and not between regriddings. Firstly, the amount of memory allocated shouldn't be increasing so much, since carpet::grid_functions remains approximately constant. Secondly, I thought that we were supposed to be using mmap for grid function data, so why do we have such a large amount of non-mmapped memory?
– SystemStatistics::hblkhd: This is the hblkhd field from mallinfo, which is supposed to be "Space allocated in mmapped regions (bytes)". This increases to about 2 GB after a couple of regriddings, but then stays roughly constant at 2 GB, which seems fine, apart from the fact that I had hoped that mmap was being used for all the gridfunction data
I suspect I got distracted doing my own investigations while I was writing it, and then lost track of it. Typically, I think I run with excess memory for each run, so that by 24 h, it hasn't grown too badly.
Further searching finds an email from someone else saying that they had memory leak problems and asking me about it. My reply was:
On 5 May 2015, at 17:14, Ian Hinder <ian.hinder@aei.mpg.de mailto:ian.hinder@aei.mpg.de> wrote:
Hi,
I have added Erik and Roland to CC, as we have been discussing this; I hope this is OK.
It sounds very similar. I am running on several hundred cores (<600) and the simulations often fail with OOM errors after less than a day. First the RSS grows, then the swap, then the OOM-killer kills it. I have observed this both on Datura and Hydra. Stopping the simulations and recovering from a checkpoint usually fixes the problem, and it runs on for another half day or so. I have done a fair amount of work on this, so I will summarise here.
*Monitor process RSS*
The process resident set size is the amount of address space which is currently mapped into physical memory by the OS. Thorn SystemStatistics can be used to measure this and put it into a Cactus variable, which can then be reduced across processes. I use:
IOBasic::outInfo_every = 1 IOBasic::outInfo_reductions = "maximum" IOBasic::outInfo_vars = " SystemStatistics::maxrss_mb SystemStatistics::swap_used_mb Carpet::gridfunctions "
and
IOScalar::outScalar_every = 128 IOScalar::outScalar_vars = " SystemStatistics::process_memory_mb Carpet::memory_procs "
SystemStatistics calls its variable "maxrss" but it should actually be called "rss", as that is what is output. maxrss is also available from the OS, and would give the maximum the RSS had ever been during the process lifetime.
Carpet::gridfunctions (in Carpet::memory_procs) measures the amount of memory Carpet has allocated in gridfunctions. For me, this remains essentially flat, whereas maxrss grows after each regridding until it reaches the maximum available, then the swap starts to grow. This indicates that the problem is not due to Carpet allocating more and more grid points due to grids changing size. It could be due to failing to free allocated memory (a leak) or freed data taking up space which cannot be used for further allocations or returned to the OS (fragmentation).
*Terminate and checkpoint on OOM*
I have a local change to SystemStatistics which adds parameters for maximum values of RSS and swap usage, above which it calls CCTK_TerminateNext, so if you have checkpoint_on_terminate, you get a clean termination and can continue the run without losing too much CPU time. I have been running with this for a couple of weeks now, and it works as advertised. I have a branch with this on, but I just realised it conflicts with a change Erik made. If you want this, let me know and I will sort it out.
*Memory profiling*
The malloc implementation in glibc provides no usable statistics. mallinfo is limited to 32 bit integers, which overflow for 64 bit systems. malloc_info, at least in the version on datura, doesn't include memory allocated via mmap. Useless. Instead, you need to use an external memory profiler to see what is going on. I have used "igprof" successfully, and this shows me that there is no "leak" of allocated memory corresponding to the increase in RSS. i.e. the problem is not caused by forgetting to free something. This suggests that the problem is fragmentation, where malloc has unallocated blocks of memory which it does not or cannot return to the OS. Malloc allocates memory in two ways: either in its main heap, or by allocating anonymous mmap regions. I had thought that only the latter could be fully returned to the OS, but this is not true. Any region of address space can be marked as unused (internally via the madvise(MADV_DONTNEED) system call) and a malloc implementation can do this on regions of its address space which have been freed. If such regions are too small (smaller than a page), then they could accumulate and not be returned to the OS.
*Alternative malloc implementations*
At the suggestion of Roland, I tried using the tcmalloc (http://gperftools.googlecode.com/git/doc/tcmalloc.html) library, which is a drop-in replacement for glibc malloc which is part of gperftools. This works fairly easily. You can compile the "minimal" version with no dependencies and then modify your optionlist:
CPPFLAGS = -I/home/rhaas/software/gperftools-2.1/include/gperftools LDFLAGS = -L/home/rhaas/software/gperftools-2.1/lib -Wl,-rpath,/home/rhaas/software/gperftools-2.1/lib -ltcmalloc_minimal
I found in one example case that this reduced the RSS process growth, so I am now using it for all my simulations. However, I still run into the same problem eventually, so it might be that it makes it better but doesn't solve it completely.
tcmalloc has an introspection interface which is presumably not as useless as glibc's malloc: http://gperftools.googlecode.com/git/doc/tcmalloc.html. I haven't tried this yet.
*Checkpoint recovery*
I noticed from the igprof profile that there are 11000 allocations (and frees) during checkpoint recovery on one process, all from the HDF5 uncompression routine. This is the "deflate" filter. When it decompresses a dataset, it allocates a buffer, initially sized the same as the compressed dataset (really dumb, as it will always need to be bigger). It then uncompresses into the buffer, "realloc"ing the buffer to twice the size each time it runs out of space. You can imagine that this might cause a lot of fragmentation. There is no tunable parameter, but we could modify the code (it's in https://svn.hdfgroup.uiuc.edu/hdf5/tags/hdf5-1_8_12/src/H5Zdeflate.c) to use a much larger starting buffer size, in the hope that this reduces the number of reallocs, and hence the amount of fragmentation. This wouldn't help the accumulated RSS, but it would probably produce a one-off decrease in the amount of fragmentation. I am currently not using periodic checkpointing, so I don't know if the compression routine has the same problem. Probably not, since it knows the output buffer size has to be smaller than the input buffer size. Apparently Frank Löffler also modified this routine, which solved some of his problems of running out of memory during recovery. Another alternative would be to disable checkpoint compression.
To see if you are suffering from the same problem, I think the quickest way would be to link against tcmalloc and use MallocExtension::instance()->GetNumericProperty(property_name, value) from tcmalloc to read off the generic.current_allocated_bytes property (Number of bytes used by the application. This will not typically match the memory use reported by the OS, because it does not include TCMalloc overhead or memory fragmentation). You could also look at the other properties they provide. Then compare this with the process RSS from systemstatistics, and Carpet's gridfunctions variable, and check to see if you actually have a memory leak, or if you are suffering from fragmentation.
This brings to mind another question: are you using HDF5 compression? If so, do you see the same problem if you switch it off?
And finally: do you get this growth on the very first segment of a simulation, or only on subsequent segments? I am thinking that checkpoint *recovery* severely fragments the heap, especially with compression, and this somehow causes growth in the RSS with each regridding.
So, while I had forgotten about all this, it turns out that I had actually thought quite a lot about it :)
-- Ian Hinder https://ianhinder.net
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 2 Aug 2018, at 08:45, Miguel Zilhão miguel.zilhao.nogueira@tecnico.ulisboa.pt wrote:
hi all,
some more information regarding this. i've ran a simulation based on the provided parfile qc0-mclachlan.par. i'm attaching a crude plot with the memory consumption, as reported by systemstatistics-process_memory and carpet-memory_procs, as function of the iteration.
the jump at iteration ~2048 corresponds to the time where some non-trivial regridding occurred. if i'm interpreting the plot correctly, while the total memory consumption of the system (as reported by systemstatistics) increases, the memory that carpet reports to be using is actually decreasing. is this a sign that something is probably leaking memory?
in case it's useful, i'm also attaching the exact parameter file i've used.
Can you do this with tcmalloc (and activate the tcmalloc thorn that I pointed you to), and plot the variables tcmalloc::
generic_current_allocated_bytes generic_heap_size tcmalloc_pageheap_free_bytes tcmalloc_pageheap_unmapped_bytes
This will let us know whether there are actual memory allocations which are being made and not freed, or whether the rss is increasing due to fragmentation.
hi Ian,
Can you do this with tcmalloc (and activate the tcmalloc thorn that I pointed you to), and plot the variables tcmalloc::
generic_current_allocated_bytes generic_heap_size tcmalloc_pageheap_free_bytes tcmalloc_pageheap_unmapped_bytes
This will let us know whether there are actual memory allocations which are being made and not freed, or whether the rss is increasing due to fragmentation.
yes, i've just re-ran with tcmalloc. please see the (very crude) plot attached. does this help? sorry, i don't really know how to interpret these data :-)
thanks! Miguel
Can you put labels on the plot for each curve? I don’t know which variables correspond to the different columns in the file.
yes, sorry. here are the labels from the tcmalloc file:
3:generic_current_allocated_bytes 4:generic_heap_size 5:tcmalloc_pageheap_free_bytes 6:tcmalloc_pageheap_unmapped_bytes
column 5 (tcmalloc_pageheap_free_bytes) is not visible in the plot since it's flat zero.
On 02/08/2018 18:41, Ian Hinder wrote:
Can you put labels on the plot for each curve? I don’t know which variables correspond to the different columns in the file.
On 2 Aug 2018, at 18:55, Miguel Zilhão miguel.zilhao.nogueira@tecnico.ulisboa.pt wrote:
yes, sorry. here are the labels from the tcmalloc file:
3:generic_current_allocated_bytes 4:generic_heap_size 5:tcmalloc_pageheap_free_bytes 6:tcmalloc_pageheap_unmapped_bytes
column 5 (tcmalloc_pageheap_free_bytes) is not visible in the plot since it's flat zero.
Hi Miguel,
It is very hard to see what is going on, because I am having to look in two different plots, with different plot ranges, and a second email with the column number definitions.
Please can you help me to help you by organising the information a bit more clearly?
It would be good to see a single plot showing all of the data, with a legend that shows what each line is. I assume that $6 in carpet-memory_procs is gridfunctions, but please check that and include the information in the plot. Also, please set the plot y axis range to start at zero; it's hard to see what's going on because the two plots have different ranges, and very confusing that one of them is not zero.
From what I can tell at the moment, the new gridfunction data takes less memory than the old after the regridding, but the rss grows dramatically. The allocated memory also grows dramatically, suggesting a memory leak, but not enough to account for the growth in rss. It's possible that something other than regridding is responsible. Do you have any HDF5 output enabled that might coincide with iteration 2048? HDF5 might be allocating buffers on first output that are then not freed because they will be used again later. Can you disable all output and see what happens?
Hello all,
it may also be good to not only look at the maximum reduction but the values on each MPI rank.
Yours, Roland
On 2 Aug 2018, at 18:55, Miguel Zilhão miguel.zilhao.nogueira@tecnico.ulisboa.pt wrote:
yes, sorry. here are the labels from the tcmalloc file:
3:generic_current_allocated_bytes 4:generic_heap_size 5:tcmalloc_pageheap_free_bytes 6:tcmalloc_pageheap_unmapped_bytes
column 5 (tcmalloc_pageheap_free_bytes) is not visible in the plot since it's flat zero.
Hi Miguel,
It is very hard to see what is going on, because I am having to look in two different plots, with different plot ranges, and a second email with the column number definitions.
Please can you help me to help you by organising the information a bit more clearly?
It would be good to see a single plot showing all of the data, with a legend that shows what each line is. I assume that $6 in carpet-memory_procs is gridfunctions, but please check that and include the information in the plot. Also, please set the plot y axis range to start at zero; it's hard to see what's going on because the two plots have different ranges, and very confusing that one of them is not zero.
From what I can tell at the moment, the new gridfunction data takes less memory than the old after the regridding, but the rss grows dramatically. The allocated memory also grows dramatically, suggesting a memory leak, but not enough to account for the growth in rss. It's possible that something other than regridding is responsible. Do you have any HDF5 output enabled that might coincide with iteration 2048? HDF5 might be allocating buffers on first output that are then not freed because they will be used again later. Can you disable all output and see what happens?
hi Ian,
sorry, i thought it was enough to see the qualitative trend of the curves... here's what i hope to be a better plot, with all the information. also, i have no hdf5 output for this run. i'm only saving plain text.
thanks, Miguel
On 02/08/2018 23:31, ian.hinder@aei.mpg.de wrote:
On 2 Aug 2018, at 18:55, Miguel Zilhão <miguel.zilhao.nogueira@tecnico.ulisboa.pt mailto:miguel.zilhao.nogueira@tecnico.ulisboa.pt> wrote:
yes, sorry. here are the labels from the tcmalloc file:
3:generic_current_allocated_bytes 4:generic_heap_size 5:tcmalloc_pageheap_free_bytes 6:tcmalloc_pageheap_unmapped_bytes
column 5 (tcmalloc_pageheap_free_bytes) is not visible in the plot since it's flat zero.
Hi Miguel,
It is very hard to see what is going on, because I am having to look in two different plots, with different plot ranges, and a second email with the column number definitions.
Please can you help me to help you by organising the information a bit more clearly?
It would be good to see a single plot showing all of the data, with a legend that shows what each line is. I assume that $6 in carpet-memory_procs is gridfunctions, but please check that and include the information in the plot. Also, please set the plot y axis range to start at zero; it's hard to see what's going on because the two plots have different ranges, and very confusing that one of them is not zero.
From what I can tell at the moment, the new gridfunction data takes less memory than the old after the regridding, but the rss grows dramatically. The allocated memory also grows dramatically, suggesting a memory leak, but not enough to account for the growth in rss. It's possible that something other than regridding is responsible. Do you have any HDF5 output enabled that might coincide with iteration 2048? HDF5 might be allocating buffers on first output that are then not freed because they will be used again later. Can you disable all output and see what happens?
-- Ian Hinder https://ianhinder.net
In addition to HDF5, traditionally CarpetIOASCII triggered memory leaks in libc. CarpetIOASCII is special in that only a single process performs I/O and thus uses much more memory than the others. You could test a run with CarpetIOASCII disabled.
-erik
On Fri, Aug 3, 2018 at 4:49 AM, Miguel Zilhão miguel.zilhao.nogueira@tecnico.ulisboa.pt wrote:
hi Ian,
sorry, i thought it was enough to see the qualitative trend of the curves... here's what i hope to be a better plot, with all the information. also, i have no hdf5 output for this run. i'm only saving plain text.
thanks, Miguel
On 02/08/2018 23:31, ian.hinder@aei.mpg.de wrote:
On 2 Aug 2018, at 18:55, Miguel Zilhão <miguel.zilhao.nogueira@tecnico.ulisboa.pt mailto:miguel.zilhao.nogueira@tecnico.ulisboa.pt> wrote:
yes, sorry. here are the labels from the tcmalloc file:
3:generic_current_allocated_bytes 4:generic_heap_size 5:tcmalloc_pageheap_free_bytes 6:tcmalloc_pageheap_unmapped_bytes
column 5 (tcmalloc_pageheap_free_bytes) is not visible in the plot since it's flat zero.
Hi Miguel,
It is very hard to see what is going on, because I am having to look in two different plots, with different plot ranges, and a second email with the column number definitions.
Please can you help me to help you by organising the information a bit more clearly?
It would be good to see a single plot showing all of the data, with a legend that shows what each line is. I assume that $6 in carpet-memory_procs is gridfunctions, but please check that and include the information in the plot. Also, please set the plot y axis range to start at zero; it's hard to see what's going on because the two plots have different ranges, and very confusing that one of them is not zero.
From what I can tell at the moment, the new gridfunction data takes less memory than the old after the regridding, but the rss grows dramatically. The allocated memory also grows dramatically, suggesting a memory leak, but not enough to account for the growth in rss. It's possible that something other than regridding is responsible. Do you have any HDF5 output enabled that might coincide with iteration 2048? HDF5 might be allocating buffers on first output that are then not freed because they will be used again later. Can you disable all output and see what happens?
-- Ian Hinder https://ianhinder.net
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
hi Erik,
i've just tried without CarpetIOASCII and i'm seeing the same pattern. the only output i have now is CarpetIOScalar (for the memory profiles) and AHFinderDirect's BH_diagnostics.ah?.gp files.
thanks, Miguel
On 03/08/2018 13:47, Erik Schnetter wrote:
In addition to HDF5, traditionally CarpetIOASCII triggered memory leaks in libc. CarpetIOASCII is special in that only a single process performs I/O and thus uses much more memory than the others. You could test a run with CarpetIOASCII disabled.
-erik
On Fri, Aug 3, 2018 at 4:49 AM, Miguel Zilhão miguel.zilhao.nogueira@tecnico.ulisboa.pt wrote:
hi Ian,
sorry, i thought it was enough to see the qualitative trend of the curves... here's what i hope to be a better plot, with all the information. also, i have no hdf5 output for this run. i'm only saving plain text.
thanks, Miguel
On 02/08/2018 23:31, ian.hinder@aei.mpg.de wrote:
On 2 Aug 2018, at 18:55, Miguel Zilhão <miguel.zilhao.nogueira@tecnico.ulisboa.pt mailto:miguel.zilhao.nogueira@tecnico.ulisboa.pt> wrote:
yes, sorry. here are the labels from the tcmalloc file:
3:generic_current_allocated_bytes 4:generic_heap_size 5:tcmalloc_pageheap_free_bytes 6:tcmalloc_pageheap_unmapped_bytes
column 5 (tcmalloc_pageheap_free_bytes) is not visible in the plot since it's flat zero.
Hi Miguel,
It is very hard to see what is going on, because I am having to look in two different plots, with different plot ranges, and a second email with the column number definitions.
Please can you help me to help you by organising the information a bit more clearly?
It would be good to see a single plot showing all of the data, with a legend that shows what each line is. I assume that $6 in carpet-memory_procs is gridfunctions, but please check that and include the information in the plot. Also, please set the plot y axis range to start at zero; it's hard to see what's going on because the two plots have different ranges, and very confusing that one of them is not zero.
From what I can tell at the moment, the new gridfunction data takes less memory than the old after the regridding, but the rss grows dramatically. The allocated memory also grows dramatically, suggesting a memory leak, but not enough to account for the growth in rss. It's possible that something other than regridding is responsible. Do you have any HDF5 output enabled that might coincide with iteration 2048? HDF5 might be allocating buffers on first output that are then not freed because they will be used again later. Can you disable all output and see what happens?
-- Ian Hinder https://ianhinder.net
Users mailing list Users@einsteintoolkit.org http://lists.einsteintoolkit.org/mailman/listinfo/users
On 3 Aug 2018, at 09:49, Miguel Zilhão miguel.zilhao.nogueira@tecnico.ulisboa.pt wrote:
hi Ian,
sorry, i thought it was enough to see the qualitative trend of the curves... here's what i hope to be a better plot, with all the information. also, i have no hdf5 output for this run. i'm only saving plain text.
Hi Miguel,
So, from this, we can see:
- Carpet is using less memory for gridfunctions after the regridding than before. I suppose this could happen if the BHs get closer to each other, and the coarser enclosed grids shrink. I'm a little surprised to see this on the very first regridding, but it's not a big effect anyway.
- The amount of *allocated* memory in the tcmalloc heap increases a bit after regridding. I don't know why this would be, since the gridfunction memory should dominate, and this decreases. It would be interesting to see this plot for longer times. If the green curve continues to step up by about 80 MB every 2048 iterations, this could be the reason for running out of memory.
- The heap size increases by much more than the additional allocated memory. This suggests that the heap has become fragmented, or that tcmalloc has not attempted to return memory to the OS. The tcmalloc thorn calls tcmalloc to release all memory to the OS in POSTREGRID, so as far as tcmalloc is concerned, no more memory can be returned. This suggests fragmentation; the free blocks are mixed up with allocated blocks so that entire pages cannot be mapped out. Can you set tcmalloc::report_every = 2048? The outputs a short summary of the heap status to stdout. I would again be interested to see whether this continues with a longer run. i.e. whether heap_size - allocated continues to increase.
- Interestingly, pageheap_unmapped grows a lot.
hi Ian,
many thanks for your analysis. for that particular run there is no significant change in the memory usage for longer times. however, i've now rerun the configuration that had originally given me problems (with lower resolution so that i'd have enough memory) with tcmalloc and i've plotted those same curves. here, the memory used increases further. i'm attaching the plot, together with the parfile and stdout output in case it's useful.
thanks, Miguel
On 04/08/2018 20:01, ian.hinder@aei.mpg.de wrote:
On 3 Aug 2018, at 09:49, Miguel Zilhão <miguel.zilhao.nogueira@tecnico.ulisboa.pt mailto:miguel.zilhao.nogueira@tecnico.ulisboa.pt> wrote:
hi Ian,
sorry, i thought it was enough to see the qualitative trend of the curves... here's what i hope to be a better plot, with all the information. also, i have no hdf5 output for this run. i'm only saving plain text.
Hi Miguel,
So, from this, we can see:
- Carpet is using less memory for gridfunctions after the regridding than before. I suppose this
could happen if the BHs get closer to each other, and the coarser enclosed grids shrink. I'm a little surprised to see this on the very first regridding, but it's not a big effect anyway.
- The amount of *allocated* memory in the tcmalloc heap increases a bit after regridding. I don't
know why this would be, since the gridfunction memory should dominate, and this decreases. It would be interesting to see this plot for longer times. If the green curve continues to step up by about 80 MB every 2048 iterations, this could be the reason for running out of memory.
- The heap size increases by much more than the additional allocated memory. This suggests that the
heap has become fragmented, or that tcmalloc has not attempted to return memory to the OS. The tcmalloc thorn calls tcmalloc to release all memory to the OS in POSTREGRID, so as far as tcmalloc is concerned, no more memory can be returned. This suggests fragmentation; the free blocks are mixed up with allocated blocks so that entire pages cannot be mapped out. Can you set tcmalloc::report_every = 2048? The outputs a short summary of the heap status to stdout. I would again be interested to see whether this continues with a longer run. i.e. whether heap_size - allocated continues to increase.
- Interestingly, pageheap_unmapped grows a lot.
-- Ian Hinder https://ianhinder.net
On 5 Aug 2018, at 15:52, Miguel Zilhão miguel.zilhao.nogueira@tecnico.ulisboa.pt wrote:
hi Ian,
many thanks for your analysis. for that particular run there is no significant change in the memory usage for longer times. however, i've now rerun the configuration that had originally given me problems (with lower resolution so that i'd have enough memory) with tcmalloc and i've plotted those same curves. here, the memory used increases further. i'm attaching the plot, together with the parfile and stdout output in case it's useful.
Hi Miguel,
The memory seems to reach a steady state by iteration ~3000. Can you run an example where it dies with an OOM?
The meaning of the different tcmalloc statistics is described at https://gperftools.github.io/gperftools/tcmalloc.html under "Generic Tcmalloc Status".
From what I see here, the amount of memory allocated by Cactus grows then reaches a steady state (the green curve). This is not accounted for in the gridfunctions, so it must be from something else. I don't know what it might be. Does this run also have HDF5 output disabled?
It looks like the allocated memory plus the unmapped memory would roughly equal the rss. That is a little surprising, since I would have expected the rss to exclude unmapped pages. Maybe until another process needs it, the kernel doesn't actually unmap it, for performance reasons (mapping it back again would cause a page fault).
Can you check this by plotting tcmalloc::generic_current_allocated + tcmalloc::pageheap_free against systemstatistics-process_memory::maxrss? If that is the case, then there is no issue with fragmentation, because even though the address space is fragmented, the "holes" have mostly been returned to the OS for other processes to use ("unmapped").
The point that Roland made also applies here: we are looking at the max across all processes and assuming that every process is the same. It's possible that one process has a high unmapped curve, but another has a high rss curve, and we don't see this on the plot. We would have to do 1D output of the grid arrays and plot each process separately to see the full detail. One way to see if this is necessary would be to plot both the max and min instead of just the max. That way, we can see if this is likely to be an issue.
hi Ian,
The memory seems to reach a steady state by iteration ~3000. Can you run an example where it dies with an OOM?
the OOM cases that i had were done in our local cluster (where i haven't compiled with tcmalloc); those were just higher resolution versions of this same simulation, where the OOM would be triggered around the time of one of these memory increases (ie, after iteration 2000 in this case, i'd guess).
From what I see here, the amount of memory allocated by Cactus grows then reaches a steady state (the green curve). This is not accounted for in the gridfunctions, so it must be from something else. I don't know what it might be. Does this run also have HDF5 output disabled?
yes, there was no HDF5 nor CarpetIOASCII output.
Can you check this by plotting tcmalloc::generic_current_allocated + tcmalloc::pageheap_free against systemstatistics-process_memory::maxrss? If that is the case, then there is no issue with fragmentation, because even though the address space is fragmented, the "holes" have mostly been returned to the OS for other processes to use ("unmapped").
sure, i've attached a plot with this.
The point that Roland made also applies here: we are looking at the max across all processes and assuming that every process is the same. It's possible that one process has a high unmapped curve, but another has a high rss curve, and we don't see this on the plot. We would have to do 1D output of the grid arrays and plot each process separately to see the full detail. One way to see if this is necessary would be to plot both the max and min instead of just the max. That way, we can see if this is likely to be an issue.
ok, i'm attaching another plot with both the min (dashed lines) and the max (full lines) plotted. i hope it helps.
thanks again, Miguel
On 8 Aug 2018, at 11:41, Miguel Zilhão miguel.zilhao.nogueira@tecnico.ulisboa.pt wrote:
hi Ian,
The memory seems to reach a steady state by iteration ~3000. Can you run an example where it dies with an OOM?
the OOM cases that i had were done in our local cluster (where i haven't compiled with tcmalloc); those were just higher resolution versions of this same simulation, where the OOM would be triggered around the time of one of these memory increases (ie, after iteration 2000 in this case, i'd guess).
Hi Miguel,
The memory problems are very likely strongly related to the machine you run on. I don't know that we can take much information from a smaller test run on a different machine. We already see from this run that Carpet is not "leaking" memory continuously; the curves for allocated memory show what has been malloced and not freed, and it remains more or less constant after the initial phase.
I think it's worth trying to get tcmalloc running on the cluster. So this means that you have never seen the OOM happen when using tcmalloc. It's possible that the improved memory allocation in tcmalloc over glibc would entirely solve the problem.
Can you check this by plotting tcmalloc::generic_current_allocated + tcmalloc::pageheap_free against systemstatistics-process_memory::maxrss? If that is the case, then there is no issue with fragmentation, because even though the address space is fragmented, the "holes" have mostly been returned to the OS for other processes to use ("unmapped").
sure, i've attached a plot with this.
Sorry, I made a mistake. It should have been pageheap_unmapped, not pageheap_free. Sorry! pageheap_free is essentially zero, and cannot account for the difference.
The point that Roland made also applies here: we are looking at the max across all processes and assuming that every process is the same. It's possible that one process has a high unmapped curve, but another has a high rss curve, and we don't see this on the plot. We would have to do 1D output of the grid arrays and plot each process separately to see the full detail. One way to see if this is necessary would be to plot both the max and min instead of just the max. That way, we can see if this is likely to be an issue.
ok, i'm attaching another plot with both the min (dashed lines) and the max (full lines) plotted. i hope it helps.
Thanks. This shows that the gridfunction usage is more or less similar across all processes, which is good. However, there is significant variation in most of the other quantities across processes. To understand this better, we would have to look at 1D ASCII output of the grid arrays, which is a bit painful to plot in gnuplot. Before this, I would definitely try to get tcmalloc running and outputting this information on the cluster in a run that actually shows the OOM. My guess is that you won't get an OOM with tcmalloc, and all will be fine :)
hi Ian,
The memory problems are very likely strongly related to the machine you run on. I don't know that we can take much information from a smaller test run on a different machine. We already see from this run that Carpet is not "leaking" memory continuously; the curves for allocated memory show what has been malloced and not freed, and it remains more or less constant after the initial phase.
I think it's worth trying to get tcmalloc running on the cluster. So this means that you have never seen the OOM happen when using tcmalloc. It's possible that the improved memory allocation in tcmalloc over glibc would entirely solve the problem.
well, i did have cases where i'd ran out of memory also in my workstation with tcmalloc (where i've been doing these tests), with this same configuration and more resolution. i don't have an OOM-killer in the workstation, though, so at some point the system would just start to swap (at which point i'd kill the job).
Sorry, I made a mistake. It should have been pageheap_unmapped, not pageheap_free. Sorry! pageheap_free is essentially zero, and cannot account for the difference.
ah, no problem. i'm attaching the updated plot.
The point that Roland made also applies here: we are looking at the max across all processes and assuming that every process is the same. It's possible that one process has a high unmapped curve, but another has a high rss curve, and we don't see this on the plot. We would have to do 1D output of the grid arrays and plot each process separately to see the full detail. One way to see if this is necessary would be to plot both the max and min instead of just the max. That way, we can see if this is likely to be an issue.
ok, i'm attaching another plot with both the min (dashed lines) and the max (full lines) plotted. i hope it helps.
Thanks. This shows that the gridfunction usage is more or less similar across all processes, which is good. However, there is significant variation in most of the other quantities across processes. To understand this better, we would have to look at 1D ASCII output of the grid arrays, which is a bit painful to plot in gnuplot. Before this, I would definitely try to get tcmalloc running and outputting this information on the cluster in a run that actually shows the OOM. My guess is that you won't get an OOM with tcmalloc, and all will be fine :)
ok, i could also try to do this on cluster once it's back online (currently it's down for maintenance).
thanks, Miguel
On 8 Aug 2018, at 12:38, Miguel Zilhão miguel.zilhao.nogueira@tecnico.ulisboa.pt wrote:
hi Ian,
The memory problems are very likely strongly related to the machine you run on. I don't know that we can take much information from a smaller test run on a different machine. We already see from this run that Carpet is not "leaking" memory continuously; the curves for allocated memory show what has been malloced and not freed, and it remains more or less constant after the initial phase. I think it's worth trying to get tcmalloc running on the cluster. So this means that you have never seen the OOM happen when using tcmalloc. It's possible that the improved memory allocation in tcmalloc over glibc would entirely solve the problem.
well, i did have cases where i'd ran out of memory also in my workstation with tcmalloc (where i've been doing these tests), with this same configuration and more resolution. i don't have an OOM-killer in the workstation, though, so at some point the system would just start to swap (at which point i'd kill the job).
OK.
Sorry, I made a mistake. It should have been pageheap_unmapped, not pageheap_free. Sorry! pageheap_free is essentially zero, and cannot account for the difference.
ah, no problem. i'm attaching the updated plot.
Good, that looks better. So we see that the rss mostly follows the sum of allocated and unmapped memory. I think one thing I have seen in the past is that a high rss is not necessarily an indication of a problem. Even thought the OS hasn't unmapped the pages from the process' address space, the memory is free if another process (or the current process) needs it. I suspect that the saturation point at iteration ~3000 is the point at which all the processes have a lot of unmapped memory, and the OS needs to start actually unmapping it, which stops the rss from growing any further.
The point that Roland made also applies here: we are looking at the max across all processes and assuming that every process is the same. It's possible that one process has a high unmapped curve, but another has a high rss curve, and we don't see this on the plot. We would have to do 1D output of the grid arrays and plot each process separately to see the full detail. One way to see if this is necessary would be to plot both the max and min instead of just the max. That way, we can see if this is likely to be an issue.
ok, i'm attaching another plot with both the min (dashed lines) and the max (full lines) plotted. i hope it helps.
Thanks. This shows that the gridfunction usage is more or less similar across all processes, which is good. However, there is significant variation in most of the other quantities across processes. To understand this better, we would have to look at 1D ASCII output of the grid arrays, which is a bit painful to plot in gnuplot. Before this, I would definitely try to get tcmalloc running and outputting this information on the cluster in a run that actually shows the OOM. My guess is that you won't get an OOM with tcmalloc, and all will be fine :)
ok, i could also try to do this on cluster once it's back online (currently it's down for maintenance).
OK. I'll be interested to see the results when you have them. The thing to look out for is generic_current_allocated growing.
users@lists.einsteintoolkit.org