A (complex) test case kills a process at different times and when it gets killed in sys_open(), depending on the system, it can that process to be become unkillable with spinning in >90% CPU in the kernel (it never leaves the kernel, or else we would be able to attach to it by ptrace).
The only option we have found to get information in that task is "echo t >/proc/sysrq-trigger":
aufs-3.18.1+2015016 produced this output for the process:
boot_int_1_1 R running 0 719 1 0x00000002
Call Trace:
[cd8e7850] [cfada850] 0xcfada850 (unreliable)
[cd8e7910] [c03972ac] __schedule+0x304/0x6b0
[cd8e7a20] [c0397a30] preempt_schedule_irq+0x4c/0x84
[cd8e7a40] [c000fb04] resume_kernel+0x84/0x94
--- interrupt: 901 at __generic_file_write_iter+0x23c/0x4a8
LR = __generic_file_write_iter+0x23c/0x4a8
[cd8e7b40] [c00be808] generic_file_write_iter+0x3c/0xec
[cd8e7b70] [c00f9db4] new_sync_write+0x84/0xdc
[cd8e7bf0] [c01c6440] xino_fwrite+0xa4/0xcc
[cd8e7c40] [c01c6498] au_xino_do_write+0x30/0x80
[cd8e7c60] [c01c6f10] au_xino_write+0x54/0xc8
[cd8e7c80] [c01d91b0] au_set_h_iptr+0xc0/0x130
[cd8e7cc0] [c01da314] au_new_inode+0x4b0/0x6a8
[cd8e7d40] [c01db0a4] aufs_lookup+0x1a8/0x298
[cd8e7d70] [c010388c] lookup_real+0x30/0x7c
[cd8e7d90] [c01089e4] do_last.isra.62+0x530/0xbc4
[cd8e7e00] [c010912c] path_openat+0xb4/0x680
[cd8e7e70] [c010a438] do_filp_open+0x30/0xa0
[cd8e7ef0] [c00f95c4] do_sys_open+0x154/0x248
[cd8e7f40] [c000f35c] ret_from_syscall+0x0/0x38
--- interrupt: c01 at 0xb7fe6ea0
LR = 0xb7fd1a28
aufs-3.18.1+-20160118 produced this slightly different trace, but this may be related to the signal hitting the process at a slightly different place during open:
boot_int_1_1 R running 0 727 712 0x00000002
Call Trace:
[cfad1880] [c0042ea4] try_to_wake_up+0xa0/0x1e8 (unreliable)
[cfad1940] [c0397efc] __schedule+0x304/0x6b0
[cfad1a50] [c0398680] preempt_schedule_irq+0x4c/0x84
[cfad1a70] [c000fb04] resume_kernel+0x84/0x94
--- interrupt: 901 at memset+0x34/0x5c
LR = new_sync_write+0x38/0xdc
[cfad1b30] [c00f9d8c] new_sync_write+0x84/0xdc (unreliable)
[cfad1bb0] [c01c6578] xino_fwrite+0xa4/0xcc
[cfad1c00] [c01c65d0] au_xino_do_write+0x30/0x80
[cfad1c20] [c01c7044] au_xino_write+0x54/0xc8
[cfad1c40] [c01d9764] au_set_h_iptr+0xc0/0x130
[cfad1c80] [c01da954] au_new_inode+0x4c0/0x6a8
[cfad1d00] [c01db368] aufs_lookup.part.38+0x188/0x29c
[cfad1d30] [c01db5f0] aufs_atomic_open+0x14c/0x370
[cfad1d90] [c0108c44] do_last.isra.62+0x7b8/0xbc4
[cfad1e00] [c0109104] path_openat+0xb4/0x680
[cfad1e70] [c010a410] do_filp_open+0x30/0xa0
[cfad1ef0] [c00f959c] do_sys_open+0x154/0x248
[cfad1f40] [c000f35c] ret_from_syscall+0x0/0x38
--- interrupt: c01 at 0xb7fe6ea0
LR = 0xb7fd1a28
The architecture used for these traces was 32-bit powerpc but we also triggered it on 64-bit powerpc and think also x86.
The aufs setup that triggers it best so far is tmpfs as upper layer, a jffs2 and 2nd layer, a loop-mounted cramfs as 3rd layer and a mtdblock-based cramfs as lower layer.
It wasn't triggered with the layers below the tmpfs being on NFS.
Hello Bernhard,
"Bernhard Kaindl":
It looks like a known (and already fixed) issue.
See the commit
5e439ff 2016-01-05 aufs: for 4.3, XINO handles EINTR from the dying process
This is highly depending upon the kernel version.
J. R. Okajima
Thanks a lot. The longterm update to 3.18.25, patch-3.18.24-25.xz, contains the upstream
296291c 2015-10-23 mm: make sendfile(2) killable
as well and the test was running on 3.18.25, so yes, the systems have been affected by it.
I can revert the 296291c on top of >=3.18.25 but if you say that the fix
5e439ff 2016-01-05 aufs: for 4.3, XINO handles EINTR from the dying process
with its parent commit would be stable for 3.18 and would be added to either a new aufs-3.18.25+ branch or the aufs-3.18.1+ branch in the next release for 3.18, I would also like go that route.
Many thanks, Bernhard
"Bernhard Kaindl":
I didn't know v3.18.25 was released and affected.
Now I know aufs needs new aufs3.18.25+ branch.
Hmm, it may take some time but I will do it in a few weeks (or months,
because I am busy now).
If you can't wait, then I'd suggest you to try backporting the commit
5e439ff 2016-01-05 aufs: for 4.3, XINO handles EINTR from the dying process
J. R. Okajima
Yes, I backported
5e439ff 2016-01-05 aufs: for 4.3, XINO handles EINTR from the dying process
and it's parent
44c72b4 aufs: tiny, extract a new func xino_fwrite_wkq()
to aufs-3.18.1+20160118 for use on linux-3.18.25+
The backport and a tiny but reliable test case is attached. Please review before use.
For me, this fixes the issue.
The only expected change that I came across was that 3.18.25 does not have vfs_writef_t (1st arg of do_xino_fwrite() and xino_fwrite_wkq()), so the diff context of do_xino_fwrite() and the 1st arg of xino_fwrite_wkq() uses au_writef_t instead.
The attached file can be applied as patch to aufs-3.18.1+20160118 but it can also extract a reliable, minimal test case and plain patch when run under /bin/bash (uses simple here-documents to create the files and echos the gcc lines to compile and to run the test case).
It is tested it with this minimal reliable and a more complex test case, but of course review it carefully and use at your own risk before use in a production environment.
"Bernhard Kaindl":
Looks good.
Thankfully I will re-use your patch when I create aufs3.18.25+.
J. R. Okajima