Hi,
We have a 3.0.51rt75 + aufs 3.0 kernel running on a dual core arm processor, and recently we have stumbled upon an occasional system hang issue that happens on multi-threaded program. On debugging we realised that this is happening because of two threads racing on read and mmap system calls and casuing a deadlock. from the comment in the aufs/f_op.c file
/*
The two i.e aufs_read/aufs_mmap follows a AB-BA locking order for si_rwsem and mmap_sem which certainly would be a problem for RT kernel as only one reader can acquire the lock. The contention on si_rwsem and mmap_sem of two threads from same process is then causing all other processes also to hang.
would it be ok to take mmap_sem lock before taking si_rwsem for all aufs IO calls in case of RT kernel? or will that be causing any other lock races in aufs?
Thanks,
Ravindra Khati
Hello Ravindra,
"Ravindra Khati":
First of all, you should know that all aufs3 series are not maintained
anymore. Aufs3.0 is unsupported since Apr 2013.
Second, you should read aufs README file and provide me the necessary
information.
But your report reminds me about a commit.
7aac34b 2014-05-06 aufs: bugfix, stop calling security_mmap_file() again
Which version of aufs are you using?
Does it have the commit?
J. R. Okajima
Hi,
Thanks for a quick reply.
We are on a quite old aufs code, aufs3.0 branch commit:f78f003f542dcade8e9fe70c1cc9e1bb39a1c8d5
and this is what I can only share from info requested in README file:
/sys/module/aufs/version: 3.0
It doesn't have the 7aac34b commit in it and it no where calls security_mmap_file(). I believe that this may be an issue with aufs4 as well, on a RT kernel (with rwsem reader restriction i.e kernels prior to 4.9) . The AB-BA locking order becomes a problem if following happens (where thread A and B are from same process)
1) Thread A: aufs_read()->si_read_lock()->acquired si_rwmem
2) Thread B: mmap(2) call -> acquired mmap_sem write lock
3) Thread A: raises page fault --> do_page_fault()(arch specific)-->acquire mmap_sem??(blocked)
4) Thread B: aufs_mmap-->si_read_lock()-->try acquiring si_rwsem??(blocked
Thanks,
Ravindra Khati
Last edit: Ravindra Khati 2018-02-26
"Ravindra Khati":
Ok, your version is so old and the commit I wrote
7aac34b 2014-05-06 aufs: bugfix, stop calling security_mmap_file() again
is unrelated.
The commit you mentioned is
f78f003 2013-03-12 Merge branch 'aufs3.0/10static' into aufs3.0/11proc_map
right?
And aufs_mmap() doesn't call security_mmap_file()?
That is strange. It should have security_mmap_file() call.
Did you change the aufs source file or apply some patches?
Since your version is unknown to me, it is difficult for me to suggest
you a solution. But the approach you wrote first seems bad.
Rather than that, I'd suggest you to unlock si_rwsem before ->mmap()
call. It should be better than yours, I guess.
J. R. Okajima
Sorry, my bad, we are certainly based on
f78f003 2013-03-12 Merge branch 'aufs3.0/10static' into aufs3.0/11proc_map
and it does have security_mmap_file(), probably I grepped a wrong string earlier.
We have some trivial patches i.e spin_lock calls changed with seq_spin_lock.
1) Pardon me for my knowledge on aufs code, I haven't digged in to it much. As you said taking mmap_sem lock before si_rwsem in IO calls would be bad: is that only from a perspective that it's going to delay the mmap calls(i.e waiting for mmap_sem lock) for some threads of a process or is there anything else?
2) Your suggestion to unlock si_rwsem before mmap(), do you mean abandon the page_fault (which is waiting to acquire the mmap_sem) raised during aufs_read and unlock the si_rwsem so that the other thread doing mmap() could acquire si_rwsem. The rwsem are implemented as recursive mutex in RT kernel, so these can be only unlocked by the thread which has locked it.
Below is the stack dumps of the threads from same process which are causing the deadlock, I hope that would give a better picture
ntp_prober D 8031a97c 0 2332 2084 0x00000000
[<8031a97c>] (__schedule+0x554/0x684) from [<8031ab48>] (schedule+0x9c/0xc4)
[<8031ab48>] (schedule+0x9c/0xc4) from [<8031bd98>] (__rt_mutex_slowlock+0x84/0xac)
[<8031bd98>] (__rt_mutex_slowlock+0x84/0xac) from [<8031c1f8>] (rt_mutex_slowlock+0xf4/0x14c)
[<8031c1f8>] (rt_mutex_slowlock+0xf4/0x14c) from [<8007c46c>] (__rt_down_read.isra.0+0x2c/0x3c)
[<8007c46c>] (__rt_down_read.isra.0+0x2c/0x3c) from [<8003b140>] (do_page_fault+0x7c/0x2bc)
[<8003b140>] (do_page_fault+0x7c/0x2bc) from [<80030324>] (do_DataAbort+0x34/0x98)
[<80030324>] (do_DataAbort+0x34/0x98) from [<80030bd0>] (__dabt_svc+0x70/0xa0)
Exception stack(0x9cce9d70 to 0x9cce9db8)
9d60: 2b556000 80805600 00000000 00000000
9d80: 9cce9e28 0000002c 00000400 00000000 80805600 80471ed0 80805600 0000002c
9da0: 00000fff 9cce9db8 800995dc 800982b0 200d0013 ffffffff
[<80030bd0>] (__dabt_svc+0x70/0xa0) from [<800982b0>] (file_read_actor+0x34/0xfc)
[<800982b0>] (file_read_actor+0x34/0xfc) from [<800995dc>] (generic_file_aio_read+0x3f0/0x6a8)
[<800995dc>] (generic_file_aio_read+0x3f0/0x6a8) from [<800bf53c>] (do_sync_read+0x98/0xd4)
[<800bf53c>] (do_sync_read+0x98/0xd4) from [<800bf9d4>] (vfs_read+0xfc/0x188)
[<800bf9d4>] (vfs_read+0xfc/0x188) from [<801895c4>] (vfsub_read_u+0xc/0x28)
[<801895c4>] (vfsub_read_u+0xc/0x28) from [<801923e0>] (aufs_read+0x8c/0x104)
[<801923e0>] (aufs_read+0x8c/0x104) from [<800bf9d4>] (vfs_read+0xfc/0x188)
[<800bf9d4>] (vfs_read+0xfc/0x188) from [<800bfda4>] (sys_read+0x34/0x68)
[<800bfda4>] (sys_read+0x34/0x68) from [<800311a0>] (ret_fast_syscall+0x0/0x30)
ntp_prober D 8031a97c 0 2333 2084 0x00000000
[<8031a97c>] (__schedule+0x554/0x684) from [<8031ab48>] (schedule+0x9c/0xc4)
[<8031ab48>] (schedule+0x9c/0xc4) from [<8031bd98>] (__rt_mutex_slowlock+0x84/0xac)
[<8031bd98>] (__rt_mutex_slowlock+0x84/0xac) from [<8031c1f8>] (rt_mutex_slowlock+0xf4/0x14c)
[<8031c1f8>] (rt_mutex_slowlock+0xf4/0x14c) from [<8007c46c>] (__rt_down_read.isra.0+0x2c/0x3c)
[<8007c46c>] (__rt_down_read.isra.0+0x2c/0x3c) from [<80181680>] (si_read_lock+0xa0/0x130)
[<80181680>] (si_read_lock+0xa0/0x130) from [<80191f50>] (aufs_mmap+0x38/0x298)
[<80191f50>] (aufs_mmap+0x38/0x298) from [<800b4224>] (mmap_region+0x218/0x418)
[<800b4224>] (mmap_region+0x218/0x418) from [<800b47d8>] (sys_mmap_pgoff+0x84/0xcc)
[<800b47d8>] (sys_mmap_pgoff+0x84/0xcc) from [<800311a0>] (ret_fast_syscall+0x0/0x30)
Thread 2332: Tries to take mmap_sem lock and goes to sleep as it's already locked by thread 2333 in mmap() call
Thread 2333: Tries to take si_rwsem lock and goes to sleep as it's already locked by thread 2332 in aufs_read().
Regards,
Ravindra Khati
Last edit: Ravindra Khati 2018-02-27
"Ravindra Khati":
si_rwsem protects au_sbinfo members. The typical case is a branch
management such as to add/del/mod branches. I am afraid that you don't
use such feature.
Do you mean read(2) systemcall issues mmap(2) internally?
I am not sure there such path exists... But I thought page-fault will
call address_space_operations instead of file_operations.
And I didn't mean disturbing page_fault. Just move forward/upward
si_read_unlock(sb) call before ->mmap() call in aufs_mmap(). As long as
other tasks don't change the member in au_sbinfo, it should be safe.
J. R. Okajima
We do have branches added which we do during boot up, but yes we don't delete or modify later and do not plan to do so.
No, I didn't mean that. What I am saying is, there are two separate threads in user space process of which, one is doing a read() system call and another is doing mmap2() system call and there is a race between these two threads for locks.
Here is the flow in kernel code for these two as they are runnig parallely.
1) Thread A: Doing read() sys call to read content of some file, here a mmapped buffer is passed which is causing a page fault: read()-->aufs_read()-->si_read_lock() lock acquired-->vfsub_read_u-->generic_file_aio_read()->page_fault-->trying to lock mmap_sem (blocked) as it's locked by "thread B"
2) Thread B: is doing mmap2() syscall to get some buffer
mmap2()->taken mmap_sem write lock(mm/mmap.c)-->aufs_mmap()-->si_read_lock()--> (blocked) as si_read_lock is acquired by thread A.
I suppose, you are reffering to h_file->f_op->mmap() call in aufs_mmap()? this isn't causing any issue here. The thread doing mmap2() sys call is blocked at the si_read_lock() in aufs_mmap, way before this internal ->mmap() call in aufs_mmap.
Regards,
Ravindra Khati
"Ravindra Khati":
"read-lock" for rwsem (aka. "shared-lock"), it is possible for TWO
thread to acquire the single rwsem, isn't it? Or your RT kernel doesn't
allow such shared-lock?
J. R. Okajima
That's the whole point, rwsem is unfortunately implemented as a recursive mutex in standard RT kernel. One thread can lock it multiple times but no two threads can have the lock simultaneously (not even read-lock :( ). This is done deliberatley in RT kernel so that the writers do not starve on readers. Since it's a typical issue of RT kernel + aufs, I believe this may be true for aufs4 as well, as I can see that locking sequence is same. In RT kernel 4.9, I can see this restriction on rwsem has been lifted off.
1) So either I need to backport that patch to our old RT kernel, I will have to see how feasible is that and will have larger impact on our system.
2) or have the mmap_sem locked in aufs_read/aufs_write before taking si_read_lock (at least in case of multithreaded programs). That certainly would make the locking order similar to mmap2() case thus avoiding deadlock. But then if I understood your earlier response correctly, this is not good from branch maintainenance pov? isn't it?
"Ravindra Khati":
Ok, now I think I got your point.
I didn't know such rwsem behaviour in RT kernel, and I thought mmap_sem
would cause a problem such like deadlock. As you wrote, if a task can
acquire mmap_sem multiple times, then it will be safe and I'd agree with
you. mmap_sem before si_rwsem is worth to try.
And if you know the detail about 4.9 RT kernel, would you tell me how to
reproduce the problem on my side? I have vanilla mainline 4.9 and latest
aufs4.9 (off course). And what do I need?
J. R. Okajima
Yes, I have tried this for the race between aufs_read and mmap, this seems to fix the issue.
but I think I would have to put this fix in all the calls where si_read_lock is being called and there are chances of page fault being raised raised in that flow.
here is the 4.9-20-rt15 kernel tar, this should have the restriction on rwsem lock
https://git.kernel.org/pub/scm/linux/kernel/git/rt/linux-stable-rt.git/commit/?h=v4.9-rt&id=702775a5e8163bb5199b917081f33a9642e72966
I got this issue by calling a getaddrinfo_a (gnu c lib, dns query function), which supposedly is creating multiple threads and causing this race to occur. But this doesn't happen every time, I just run a loop in shell to run the binary multiple times. I am also trying to reproduce this with a simple multithreaded program, doing mmap and read calls in each threads but so far I am not able to reproduce it this way.
Ravindra Khati