Hi,
I have added Erik and Roland to CC, as we have been discussing this; I hope this is OK.
It sounds very similar. I am running on several hundred cores (<600) and the simulations often fail with OOM errors after less than a day. First the RSS grows, then the swap, then the OOM-killer kills it. I have observed this both on Datura and Hydra. Stopping the simulations and recovering from a checkpoint usually fixes the problem, and it runs on for another half day or so. I have done a fair amount of work on this, so I will summarise here.
Monitor process RSS
The process resident set size is the amount of address space which is currently mapped into physical memory by the OS. Thorn SystemStatistics can be used to measure this and put it into a Cactus variable, which can then be reduced across processes. I use:
IOBasic::outInfo_every = 1
IOBasic::outInfo_reductions = "maximum"
IOBasic::outInfo_vars = "
SystemStatistics::maxrss_mb
SystemStatistics::swap_used_mb
Carpet::gridfunctions
"
and
IOScalar::outScalar_every = 128
IOScalar::outScalar_vars = "
SystemStatistics::process_memory_mb
Carpet::memory_procs
"
SystemStatistics calls its variable "maxrss" but it should actually be called "rss", as that is what is output. maxrss is also available from the OS, and would give the maximum the RSS had ever been during the process lifetime.
Carpet::gridfunctions (in Carpet::memory_procs) measures the amount of memory Carpet has allocated in gridfunctions. For me, this remains essentially flat, whereas maxrss grows after each regridding until it reaches the maximum available, then the swap starts to grow. This indicates that the problem is not due to Carpet allocating more and more grid points due to grids changing size. It could be due to failing to free allocated memory (a leak) or freed data taking up space which cannot be used for further allocations or returned to the OS (fragmentation).
Terminate and checkpoint on OOM
I have a local change to SystemStatistics which adds parameters for maximum values of RSS and swap usage, above which it calls CCTK_TerminateNext, so if you have checkpoint_on_terminate, you get a clean termination and can continue the run without losing too much CPU time. I have been running with this for a couple of weeks now, and it works as advertised. I have a branch with this on, but I just realised it conflicts with a change Erik made. If you want this, let me know and I will sort it out.
Memory profiling
The malloc implementation in glibc provides no usable statistics. mallinfo is limited to 32 bit integers, which overflow for 64 bit systems. malloc_info, at least in the version on datura, doesn't include memory allocated via mmap. Useless. Instead, you need to use an external memory profiler to see what is going on. I have used "igprof" successfully, and this shows me that there is no "leak" of allocated memory corresponding to the increase in RSS. i.e. the problem is not caused by forgetting to free something. This suggests that the problem is fragmentation, where malloc has unallocated blocks of memory which it does not or cannot return to the OS. Malloc allocates memory in two ways: either in its main heap, or by allocating anonymous mmap regions. I had thought that only the latter could be fully returned to the OS, but this is not true. Any region of address space can be marked as unused (internally via the madvise(MADV_DONTNEED) system call) and a malloc implementation can do this on regions of its address space which have been freed. If such regions are too small (smaller than a page), then they could accumulate and not be returned to the OS.
Alternative malloc implementations
At the suggestion of Roland, I tried using the tcmalloc (
http://gperftools.googlecode.com/git/doc/tcmalloc.html) library, which is a drop-in replacement for glibc malloc which is part of gperftools. This works fairly easily. You can compile the "minimal" version with no dependencies and then modify your optionlist:
CPPFLAGS = -I/home/rhaas/software/gperftools-2.1/include/gperftools
LDFLAGS = -L/home/rhaas/software/gperftools-2.1/lib -Wl,-rpath,/home/rhaas/software/gperftools-2.1/lib -ltcmalloc_minimal
I found in one example case that this reduced the RSS process growth, so I am now using it for all my simulations. However, I still run into the same problem eventually, so it might be that it makes it better but doesn't solve it completely.
Checkpoint recovery
I noticed from the igprof profile that there are 11000 allocations (and frees) during checkpoint recovery on one process, all from the HDF5 uncompression routine. This is the "deflate" filter. When it decompresses a dataset, it allocates a buffer, initially sized the same as the compressed dataset (really dumb, as it will always need to be bigger). It then uncompresses into the buffer, "realloc"ing the buffer to twice the size each time it runs out of space. You can imagine that this might cause a lot of fragmentation. There is no tunable parameter, but we could modify the code (it's in
https://svn.hdfgroup.uiuc.edu/hdf5/tags/hdf5-1_8_12/src/H5Zdeflate.c) to use a much larger starting buffer size, in the hope that this reduces the number of reallocs, and hence the amount of fragmentation. This wouldn't help the accumulated RSS, but it would probably produce a one-off decrease in the amount of fragmentation. I am currently not using periodic checkpointing, so I don't know if the compression routine has the same problem. Probably not, since it knows the output buffer size has to be smaller than the input buffer size. Apparently Frank Löffler also modified this routine, which solved some of his problems of running out of memory during recovery. Another alternative would be to disable checkpoint compression.
To see if you are suffering from the same problem, I think the quickest way would be to link against tcmalloc and use MallocExtension::instance()->GetNumericProperty(property_name, value) from tcmalloc to read off the generic.current_allocated_bytes property (Number of bytes used by the application. This will not typically match the memory use reported by the OS, because it does not include TCMalloc overhead or memory fragmentation). You could also look at the other properties they provide. Then compare this with the process RSS from systemstatistics, and Carpet's gridfunctions variable, and check to see if you actually have a memory leak, or if you are suffering from fragmentation.