kernel-ark

History

Roman Pen 68ac546f26 mm/vmalloc: fix possible exhaustion of vmalloc space caused by vm_map_ram allocator Recently I came across high fragmentation of vm_map_ram allocator: vmap_block has free space, but still new blocks continue to appear. Further investigation showed that certain mapping/unmapping sequences can exhaust vmalloc space. On small 32bit systems that's not a big problem, cause purging will be called soon on a first allocation failure (alloc_vmap_area), but on 64bit machines, e.g. x86_64 has 45 bits of vmalloc space, that can be a disaster. 1) I came up with a simple allocation sequence, which exhausts virtual space very quickly: while (iters) { /* Map/unmap big chunk / vaddr = vm_map_ram(pages, 16, -1, PAGE_KERNEL); vm_unmap_ram(vaddr, 16); / Map/unmap small chunks. * * -1 for hole, which should be left at the end of each block * to keep it partially used, with some free space available / for (i = 0; i < (VMAP_BBMAP_BITS - 16) / 8 - 1; i++) { vaddr = vm_map_ram(pages, 8, -1, PAGE_KERNEL); vm_unmap_ram(vaddr, 8); } } The idea behind is simple: 1. We have to map a big chunk, e.g. 16 pages. 2. Then we have to occupy the remaining space with smaller chunks, i.e. 8 pages. At the end small hole should remain to keep block in free list, but do not let big chunk to occupy remaining space. 3. Goto 1 - allocation request of 16 pages can't be completed (only 8 slots are left free in the block in the #2 step), new block will be allocated, all further requests will lay into newly allocated block. To have some measurement numbers for all further tests I setup ftrace and enabled 4 basic calls in a function profile: echo vm_map_ram > /sys/kernel/debug/tracing/set_ftrace_filter; echo alloc_vmap_area >> /sys/kernel/debug/tracing/set_ftrace_filter; echo vm_unmap_ram >> /sys/kernel/debug/tracing/set_ftrace_filter; echo free_vmap_block >> /sys/kernel/debug/tracing/set_ftrace_filter; So for this scenario I got these results: BEFORE (all new blocks are put to the head of a free list) # cat /sys/kernel/debug/tracing/trace_stat/function0 Function Hit Time Avg s^2 -------- --- ---- --- --- vm_map_ram 126000 30683.30 us 0.243 us 30819.36 us vm_unmap_ram 126000 22003.24 us 0.174 us 340.886 us alloc_vmap_area 1000 4132.065 us 4.132 us 0.903 us AFTER (all new blocks are put to the tail of a free list) # cat /sys/kernel/debug/tracing/trace_stat/function0 Function Hit Time Avg s^2 -------- --- ---- --- --- vm_map_ram 126000 28713.13 us 0.227 us 24944.70 us vm_unmap_ram 126000 20403.96 us 0.161 us 1429.872 us alloc_vmap_area 993 3916.795 us 3.944 us 29.370 us free_vmap_block 992 654.157 us 0.659 us 1.273 us SUMMARY: The most interesting numbers in those tables are numbers of block allocations and deallocations: alloc_vmap_area and free_vmap_block calls, which show that before the change blocks were not freed, and virtual space and physical memory (vmap_block structure allocations, etc) were consumed. Average time which were spent in vm_map_ram/vm_unmap_ram became slightly better. That can be explained with a reasonable amount of blocks in a free list, which we need to iterate to find a suitable free block. 2) Another scenario is a random allocation: while (iters) { / Randomly take number from a range [1..32/64] / nr = rand(1, VMAP_MAX_ALLOC); vaddr = vm_map_ram(pages, nr, -1, PAGE_KERNEL); vm_unmap_ram(vaddr, nr); } I chose mersenne twister PRNG to generate persistent random state to guarantee that both runs have the same random sequence. For each vm_map_ram call random number from [1..32/64] was taken to represent amount of pages which I do map. I did 10'000 vm_map_ram calls and got these two tables: BEFORE (all new blocks are put to the head of a free list) # cat /sys/kernel/debug/tracing/trace_stat/function0 Function Hit Time Avg s^2 -------- --- ---- --- --- vm_map_ram 10000 10170.01 us 1.017 us 993.609 us vm_unmap_ram 10000 5321.823 us 0.532 us 59.789 us alloc_vmap_area 420 2150.239 us 5.119 us 3.307 us free_vmap_block 37 159.587 us 4.313 us 134.344 us AFTER (all new blocks are put to the tail of a free list) # cat /sys/kernel/debug/tracing/trace_stat/function0 Function Hit Time Avg s^2 -------- --- ---- --- --- vm_map_ram 10000 7745.637 us 0.774 us 395.229 us vm_unmap_ram 10000 5460.573 us 0.546 us 67.187 us alloc_vmap_area 414 2201.650 us 5.317 us 5.591 us free_vmap_block 412 574.421 us 1.394 us 15.138 us SUMMARY: 'BEFORE' table shows, that 420 blocks were allocated and only 37 were freed. Remained 383 blocks are still in a free list, consuming virtual space and physical memory. 'AFTER' table shows, that 414 blocks were allocated and 412 were really freed. 2 blocks remained in a free list. So fragmentation was dramatically reduced. Why? Because when we put newly allocated block to the head, all further requests will occupy new block, regardless remained space in other blocks. In this scenario all requests come randomly. Eventually remained free space will be less than requested size, free list will be iterated and it is possible that nothing will be found there - finally new block will be created. So exhaustion in random scenario happens for the maximum possible allocation size: 32 pages for 32-bit system and 64 pages for 64-bit system. Also average cost of vm_map_ram was reduced from 1.017 us to 0.774 us. Again this can be explained by iteration through smaller list of free blocks. 3) Next simple scenario is a sequential allocation, when the allocation order is increased for each block. This scenario forces allocator to reach maximum amount of partially free blocks in a free list: while (iters) { / Populate free list with blocks with remaining space / for (order = 0; order <= ilog2(VMAP_MAX_ALLOC); order++) { nr = VMAP_BBMAP_BITS / (1 << order); / Leave a hole / nr -= 1; for (i = 0; i < nr; i++) { vaddr = vm_map_ram(pages, (1 << order), -1, PAGE_KERNEL); vm_unmap_ram(vaddr, (1 << order)); } / Completely occupy blocks from a free list / for (order = 0; order <= ilog2(VMAP_MAX_ALLOC); order++) { vaddr = vm_map_ram(pages, (1 << order), -1, PAGE_KERNEL); vm_unmap_ram(vaddr, (1 << order)); } } Results which I got: BEFORE (all new blocks are put to the head of a free list) # cat /sys/kernel/debug/tracing/trace_stat/function0 Function Hit Time Avg s^2 -------- --- ---- --- --- vm_map_ram 2032000 399545.2 us 0.196 us 467123.7 us vm_unmap_ram 2032000 363225.7 us 0.178 us 111405.9 us alloc_vmap_area 7001 30627.76 us 4.374 us 495.755 us free_vmap_block 6993 7011.685 us 1.002 us 159.090 us AFTER (all new blocks are put to the tail of a free list) # cat /sys/kernel/debug/tracing/trace_stat/function0 Function Hit Time Avg s^2 -------- --- ---- --- --- vm_map_ram 2032000 394259.7 us 0.194 us 589395.9 us vm_unmap_ram 2032000 292500.7 us 0.143 us 94181.08 us alloc_vmap_area 7000 31103.11 us 4.443 us 703.225 us free_vmap_block 7000 6750.844 us 0.964 us 119.112 us SUMMARY: No surprises here, almost all numbers are the same. Fixing this fragmentation problem I also did some improvements in a allocation logic of a new vmap block: occupy block immediately and get rid of extra search in a free list. Also I replaced dirty bitmap with min/max dirty range values to make the logic simpler and slightly faster, since two longs comparison costs less, than loop thru bitmap. This patchset raises several questions: Q: Think the problem you comments is already known so that I wrote comments about it as "it could consume lots of address space through fragmentation". Could you tell me about your situation and reason why it should be avoided? Gioh Kim A: Indeed, there was a commit `364376383` which adds explicit comment about fragmentation. But fragmentation which is described in this comment caused by mixing of long-lived and short-lived objects, when a whole block is pinned in memory because some page slots are still in use. But here I am talking about blocks which are free, nobody uses them, and allocator keeps them alive forever, continuously allocating new blocks. Q: I think that if you put newly allocated block to the tail of a free list, below example would results in enormous performance degradation. new block: 1MB (256 pages) while (iters--) { vm_map_ram(3 or something else not dividable for 256) 85 vm_unmap_ram(3) * 85 } On every iteration, it needs newly allocated block and it is put to the tail of a free list so finding it consumes large amount of time. Joonsoo Kim A: Second patch in current patchset gets rid of extra search in a free list, so new block will be immediately occupied.. Also, the scenario above is impossible, cause vm_map_ram allocates virtual range in orders, i.e. 2^n. I.e. passing 3 to vm_map_ram you will allocate 4 slots in a block and 256 slots (capacity of a block) of course dividable on 4, so block will be completely occupied. But there is a worst case which we can achieve: each free block has a hole equal to order size. The maximum size of allocation is 64 pages for 64-bit system (if you try to map more, original alloc_vmap_area will be called). So the maximum order is 6. That means that worst case, before allocator makes a decision to allocate a new block, is to iterate 7 blocks: HEAD 1st block - has 1 page slot free (order 0) 2nd block - has 2 page slots free (order 1) 3rd block - has 4 page slots free (order 2) 4th block - has 8 page slots free (order 3) 5th block - has 16 page slots free (order 4) 6th block - has 32 page slots free (order 5) 7th block - has 64 page slots free (order 6) TAIL So the worst scenario on 64-bit system is that each CPU queue can have 7 blocks in a free list. This can happen only and only if you allocate blocks increasing the order. (as I did in the function written in the comment of the first patch) This is weird and rare case, but still it is possible. Afterwards you will get 7 blocks in a list. All further requests should be placed in a newly allocated block or some free slots should be found in a free list. Seems it does not look dramatically awful. This patch (of 3): If suitable block can't be found, new block is allocated and put into a head of a free list, so on next iteration this new block will be found first. That's bad, because old blocks in a free list will not get a chance to be fully used, thus fragmentation will grow. Let's consider this simple example: #1 We have one block in a free list which is partially used, and where only one page is free: HEAD \|xxxxxxxxx-\| TAIL ^ free space for 1 page, order 0 #2 New allocation request of order 1 (2 pages) comes, new block is allocated since we do not have free space to complete this request. New block is put into a head of a free list: HEAD \|----------\|xxxxxxxxx-\| TAIL #3 Two pages were occupied in a new found block: HEAD \|xx--------\|xxxxxxxxx-\| TAIL ^ two pages mapped here #4 New allocation request of order 0 (1 page) comes. Block, which was created on #2 step, is located at the beginning of a free list, so it will be found first: HEAD \|xxX-------\|xxxxxxxxx-\| TAIL ^ ^ page mapped here, but better to use this hole It is obvious, that it is better to complete request of #4 step using the old block, where free space is left, because in other case fragmentation will be highly increased. But fragmentation is not only the case. The worst thing is that I can easily create scenario, when the whole vmalloc space is exhausted by blocks, which are not used, but already dirty and have several free pages. Let's consider this function which execution should be pinned to one CPU: static void exhaust_virtual_space(struct page pages[16], int iters) { / Firstly we have to map a big chunk, e.g. 16 pages. * Then we have to occupy the remaining space with smaller * chunks, i.e. 8 pages. At the end small hole should remain. * So at the end of our allocation sequence block looks like * this: * XX big chunk * \|XXxxxxxxx-\| x small chunk * - hole, which is enough for a small chunk, * but is not enough for a big chunk / while (iters--) { int i; void vaddr; /* Map/unmap big chunk / vaddr = vm_map_ram(pages, 16, -1, PAGE_KERNEL); vm_unmap_ram(vaddr, 16); / Map/unmap small chunks. * * -1 for hole, which should be left at the end of each block * to keep it partially used, with some free space available */ for (i = 0; i < (VMAP_BBMAP_BITS - 16) / 8 - 1; i++) { vaddr = vm_map_ram(pages, 8, -1, PAGE_KERNEL); vm_unmap_ram(vaddr, 8); } } } On every iteration new block (1MB of vm area in my case) will be allocated and then will be occupied, without attempt to resolve small allocation request using previously allocated blocks in a free list. In case of random allocation (size should be randomly taken from the range [1..64] in 64-bit case or [1..32] in 32-bit case) situation is the same: new blocks continue to appear if maximum possible allocation size (32 or 64) passed to the allocator, because all remaining blocks in a free list do not have enough free space to complete this allocation request. In summary if new blocks are put into the head of a free list eventually virtual space will be exhausted. In current patch I simply put newly allocated block to the tail of a free list, thus reduce fragmentation, giving a chance to resolve allocation request using older blocks with possible holes left. Signed-off-by: Roman Pen <r.peniaev@gmail.com> Cc: Eric Dumazet <edumazet@google.com> Acked-by: Joonsoo Kim <iamjoonsoo.kim@lge.com> Cc: David Rientjes <rientjes@google.com> Cc: WANG Chao <chaowang@redhat.com> Cc: Fabian Frederick <fabf@skynet.be> Cc: Christoph Lameter <cl@linux.com> Cc: Gioh Kim <gioh.kim@lge.com> Cc: Rob Jones <rob.jones@codethink.co.uk> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>		2015-04-15 16:35:18 -07:00
..
kasan	kasan, module, vmalloc: rework shadow allocation for modules	2015-03-12 18:46:08 -07:00
backing-dev.c	Merge branch 'lazytime' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs	2015-02-17 16:12:34 -08:00
balloon_compaction.c	mm/balloon_compaction: fix deflation when compaction is disabled	2014-10-29 16:33:15 -07:00
bootmem.c	mem-hotplug: reset node managed pages when hot-adding a new pgdat	2014-11-13 16:17:06 -08:00
cleancache.c	cleancache: remove limit on the number of cleancache enabled filesystems	2015-04-14 16:49:03 -07:00
cma_debug.c	mm-cma-allocation-trigger-fix	2015-04-14 16:49:00 -07:00
cma.c	mm: cma: constify and use correct signness in mm/cma.c	2015-04-14 16:49:04 -07:00
cma.h	mm: cma: allocation trigger	2015-04-14 16:49:00 -07:00
compaction.c	mm/compaction: reset compaction scanner positions	2015-04-15 16:35:17 -07:00
debug-pagealloc.c	mm/debug-pagealloc: make debug-pagealloc boottime configurable	2014-12-13 12:42:48 -08:00
debug.c	mm: account pmd page tables to the process	2015-02-11 17:06:04 -08:00
dmapool.c	mm/dmapool.c: fixed a brace coding style issue	2014-10-09 22:26:00 -04:00
early_ioremap.c	mm: create generic early_ioremap() support	2014-04-07 16:36:15 -07:00
fadvise.c	vfs: remove get_xip_mem	2015-02-16 17:56:03 -08:00
failslab.c
filemap.c	Merge branch 'akpm' (patches from Andrew)	2015-04-14 16:49:17 -07:00
frontswap.c	mm/frontswap.c: fix the condition in BUG_ON	2014-12-10 17:41:08 -08:00
gup.c	mm: move mm_populate()-related code to mm/gup.c	2015-04-14 16:49:00 -07:00
highmem.c	mm/highmem: make kmap cache coloring aware	2014-08-06 18:01:22 -07:00
huge_memory.c	mm, memcg: sync allocation and memcg charge gfp flags for THP	2015-04-15 16:35:17 -07:00
hugetlb_cgroup.c	mm: page_counter: pull "-1" handling out of page_counter_memparse()	2015-02-11 17:06:02 -08:00
hugetlb.c	hugetlbfs: accept subpool min_size mount option and setup accordingly	2015-04-15 16:35:18 -07:00
hwpoison-inject.c	mm/hwpoison-inject.c: remove unnecessary null test before debugfs_remove_recursive	2014-08-06 18:01:19 -07:00
init-mm.c
internal.h	mm/compaction: enhance compaction finish condition	2015-04-14 16:49:01 -07:00
interval_tree.c	mm: replace vma->sharead.linear with vma->shared	2015-02-10 14:30:31 -08:00
Kconfig	mm: cma: debugfs interface	2015-04-14 16:49:00 -07:00
Kconfig.debug	mm/debug_pagealloc: remove obsolete Kconfig options	2015-01-08 15:10:52 -08:00
kmemcheck.c	mm/slab_common: move kmem_cache definition to internal header	2014-10-09 22:25:50 -04:00
kmemleak-test.c	mm/kmemleak-test.c: use pr_fmt for logging	2014-06-06 16:08:18 -07:00
kmemleak.c	kmemleak: disable kasan instrumentation for kmemleak	2015-02-13 21:21:41 -08:00
ksm.c	mm: remove rest usage of VM_NONLINEAR and pte_file()	2015-02-10 14:30:31 -08:00
list_lru.c	memcg: reparent list_lrus and free kmemcg_id on css offline	2015-02-12 18:54:10 -08:00
maccess.c
madvise.c	vfs: remove get_xip_mem	2015-02-16 17:56:03 -08:00
Makefile	mm: move memtest under mm	2015-04-14 16:49:06 -07:00
memblock.c	mm/memblock.c: rename local variable of memblock_type to `type'	2015-04-14 16:49:00 -07:00
memcontrol.c	memcg: remove obsolete comment	2015-04-15 16:35:16 -07:00
memory_hotplug.c	mm, hotplug: fix concurrent memory hot-add deadlock	2015-04-14 16:49:00 -07:00
memory-failure.c	mm/memory-failure.c: define page types for action_result() in one place	2015-04-15 16:35:16 -07:00
memory.c	mm: refactor do_wp_page handling of shared vma into a function	2015-04-14 16:49:03 -07:00
mempolicy.c	mm, thp: really limit transparent hugepage allocation to local node	2015-04-14 16:49:03 -07:00
mempool.c	mm, mempool: do not allow atomic resizing	2015-04-14 16:49:06 -07:00
memtest.c	memtest: use phys_addr_t for physical addresses	2015-04-14 16:49:06 -07:00
migrate.c	mm/migrate: check-before-clear PageSwapCache	2015-04-15 16:35:17 -07:00
mincore.c	mincore: apply page table walker on do_mincore()	2015-02-11 17:06:06 -08:00
mlock.c	mm: move mm_populate()-related code to mm/gup.c	2015-04-14 16:49:00 -07:00
mm_init.c	mm/mm_init.c: mark mminit_loglevel __meminitdata	2015-02-12 18:54:11 -08:00
mmap.c	mm: rename __mlock_vma_pages_range() to populate_vma_page_range()	2015-04-14 16:49:00 -07:00
mmu_context.c	sched/mm: call finish_arch_post_lock_switch in idle_task_exit and use_mm	2014-02-21 08:50:17 +01:00
mmu_notifier.c	mmu_notifier: add the callback for mmu_notifier_invalidate_range()	2014-11-13 13:46:09 +11:00
mmzone.c	mm: microoptimize zonelist operations	2015-02-11 17:06:02 -08:00
mprotect.c	mm: numa: preserve PTE write permissions across a NUMA hinting fault	2015-03-25 16:20:31 -07:00
mremap.c	fix mremap() vs. ioctx_kill() race	2015-04-06 17:50:59 -04:00
msync.c	mm: remove rest usage of VM_NONLINEAR and pte_file()	2015-02-10 14:30:31 -08:00
nobootmem.c	mem-hotplug: reset node managed pages when hot-adding a new pgdat	2014-11-13 16:17:06 -08:00
nommu.c	mm/nommu.c: export symbol max_mapnr	2015-03-12 18:46:08 -07:00
oom_kill.c	mm/oom_kill.c: fix typo in comment	2015-04-15 16:35:16 -07:00
page_alloc.c	mm/page_alloc.c: clean up comment	2015-04-14 16:49:04 -07:00
page_counter.c	mm: page_counter: pull "-1" handling out of page_counter_memparse()	2015-02-11 17:06:02 -08:00
page_ext.c	mm/page_owner: keep track of page owners	2014-12-13 12:42:48 -08:00
page_io.c	fs: move struct kiocb to fs.h	2015-03-25 20:28:11 -04:00
page_isolation.c	mm/page_alloc.c: call kernel_map_pages in unset_migrateype_isolate	2015-03-25 16:20:30 -07:00
page_owner.c	mm/page_owner.c: remove unnecessary stack_trace field	2015-02-11 17:06:07 -08:00
page-writeback.c	mm/page-writeback: check-before-clear PageReclaim	2015-04-15 16:35:17 -07:00
pagewalk.c	mm/pagewalk.c: prevent positive return value of walk_page_test() from being passed to callers	2015-03-25 16:20:30 -07:00
percpu-km.c	percpu: implmeent pcpu_nr_empty_pop_pages and chunk->nr_populated	2014-09-02 14:46:05 -04:00
percpu-vm.c	percpu: move region iterations out of pcpu_[de]populate_chunk()	2014-09-02 14:46:02 -04:00
percpu.c	percpu: Fix trivial typos in comments	2015-03-24 13:41:54 -04:00
pgtable-generic.c	mm: convert p[te\|md]_mknonnuma and remaining page table manipulations	2015-02-12 18:54:08 -08:00
process_vm_access.c	process_vm_access: switch to {compat_,}import_iovec()	2015-04-11 22:27:12 -04:00
quicklist.c
readahead.c	fs: export inode_to_bdi and use it in favor of mapping->backing_dev_info	2015-01-20 14:03:04 -07:00
rmap.c	mm: fix anon_vma->degree underflow in anon_vma endless growing prevention	2015-03-25 16:20:30 -07:00
shmem.c	Merge branch 'iocb' into for-next	2015-04-11 22:24:41 -04:00
slab_common.c	mm: slub: add kernel address sanitizer support for slub allocator	2015-02-13 21:21:41 -08:00
slab.c	mm: remove GFP_THISNODE	2015-04-14 16:49:03 -07:00
slab.h	slub: make dead caches discard free slabs immediately	2015-02-12 18:54:10 -08:00
slob.c	slob: make slob_alloc_node() static and remove EXPORT_SYMBOL()	2015-04-14 16:48:59 -07:00
slub.c	slub: use bool function return values of true/false not 1/0	2015-04-14 16:48:59 -07:00
sparse-vmemmap.c
sparse.c	mm: use macros from compiler.h instead of __attribute__((...))	2014-04-07 16:35:54 -07:00
swap_cgroup.c	mm: page_cgroup: rename file to mm/swap_cgroup.c	2014-12-10 17:41:09 -08:00
swap_state.c	fs: remove mapping->backing_dev_info	2015-01-20 14:03:05 -07:00
swap.c	mm: rename deactivate_page to deactivate_file_page	2015-04-15 16:35:17 -07:00
swapfile.c	mm: page_cgroup: rename file to mm/swap_cgroup.c	2014-12-10 17:41:09 -08:00
truncate.c	mm: rename deactivate_page to deactivate_file_page	2015-04-15 16:35:17 -07:00
util.c	mm/util: add kstrdup_const	2015-02-13 21:21:35 -08:00
vmacache.c	mm,vmacache: count number of system-wide flushes	2014-12-13 12:42:48 -08:00
vmalloc.c	mm/vmalloc: fix possible exhaustion of vmalloc space caused by vm_map_ram allocator	2015-04-15 16:35:18 -07:00
vmpressure.c	mm/vmpressure.c: fix race in vmpressure_work_fn()	2014-12-02 17:32:07 -08:00
vmscan.c	Merge branch 'akpm' (patches from Andrew)	2015-02-12 18:54:28 -08:00
vmstat.c	vmstat: Reduce time interval to stat update on idle cpu	2015-02-11 17:06:07 -08:00
workingset.c	list_lru: add helpers to isolate items	2015-02-12 18:54:10 -08:00
zbud.c	mm/zpool: add name argument to create zpool	2015-02-12 18:54:12 -08:00
zpool.c	mm/zpool: add name argument to create zpool	2015-02-12 18:54:12 -08:00
zsmalloc.c	mm/zsmalloc: add statistics support	2015-02-12 18:54:12 -08:00
zswap.c	mm/zpool: add name argument to create zpool	2015-02-12 18:54:12 -08:00