Menu ▾ ▴

#28 RT System (deadlocks) hangs because of locking order in case of aufs_read/aufs_mmap.

v1.0_(example)
open
nobody
None
3
2018-03-01
2018-02-26
No

Hi,

We have a 3.0.51rt75 + aufs 3.0 kernel running on a dual core arm processor, and recently we have stumbled upon an occasional system hang issue that happens on multi-threaded program. On debugging we realised that this is happening because of two threads racing on read and mmap system calls and casuing a deadlock. from the comment in the aufs/f_op.c file

/*

  • The locking order around current->mmap_sem.
    • in most and regular cases
  • file I/O syscall -- aufs_read() or something
  • -- si_rwsem for read -- mmap_sem
  • (Note that [fdi]i_rwsem are released before mmap_sem).
    • in mmap case
  • mmap(2) -- mmap_sem -- aufs_mmap() -- si_rwsem for read -- [fdi]i_rwsem
  • This AB-BA order is definitly bad, but is not a problem since "si_rwsem for
  • read" allows muliple processes to acquire it and [fdi]i_rwsem are not held in
  • file I/O.*/

The two i.e aufs_read/aufs_mmap follows a AB-BA locking order for si_rwsem and mmap_sem which certainly would be a problem for RT kernel as only one reader can acquire the lock. The contention on si_rwsem and mmap_sem of two threads from same process is then causing all other processes also to hang.

would it be ok to take mmap_sem lock before taking si_rwsem for all aufs IO calls in case of RT kernel? or will that be causing any other lock races in aufs?

Thanks,
Ravindra Khati

Discussion

  • J. R. Okajima

    J. R. Okajima - 2018-02-26

    Hello Ravindra,

    "Ravindra Khati":

    We have a 3.0.51rt75 + aufs 3.0 kernel running on a dual core arm processor, and recently we have stumbled upon an occasional system hang issue that happens on multi-threaded program. On debugging we realised that this is happening because of two threads racing on read and mmap system calls and casuing a deadlock. from the comment in the aufs/f_op.c file

    First of all, you should know that all aufs3 series are not maintained
    anymore. Aufs3.0 is unsupported since Apr 2013.
    Second, you should read aufs README file and provide me the necessary
    information.

    But your report reminds me about a commit.
    7aac34b 2014-05-06 aufs: bugfix, stop calling security_mmap_file() again

    Which version of aufs are you using?
    Does it have the commit?

    J. R. Okajima

     
  • Ravindra Khati

    Ravindra Khati - 2018-02-26

    Hi,

    Thanks for a quick reply.
    We are on a quite old aufs code, aufs3.0 branch commit:f78f003f542dcade8e9fe70c1cc9e1bb39a1c8d5

    and this is what I can only share from info requested in README file:
    /sys/module/aufs/version: 3.0

    It doesn't have the 7aac34b commit in it and it no where calls security_mmap_file(). I believe that this may be an issue with aufs4 as well, on a RT kernel (with rwsem reader restriction i.e kernels prior to 4.9) . The AB-BA locking order becomes a problem if following happens (where thread A and B are from same process)
    1) Thread A: aufs_read()->si_read_lock()->acquired si_rwmem
    2) Thread B: mmap(2) call -> acquired mmap_sem write lock
    3) Thread A: raises page fault --> do_page_fault()(arch specific)-->acquire mmap_sem??(blocked)
    4) Thread B: aufs_mmap-->si_read_lock()-->try acquiring si_rwsem??(blocked

    Thanks,
    Ravindra Khati

     

    Last edit: Ravindra Khati 2018-02-26
    • J. R. Okajima

      J. R. Okajima - 2018-02-27

      "Ravindra Khati":

      We are on a quite old aufs code, aufs3.0 branch commit:f78f003f542dcade8e9fe70c1cc9e1bb39a1c8d5

      and this is what I can only share from info requested in README file:
      /sys/module/aufs/version: 3.0

      It doesn't have the 7aac34b commit in it and it no where calls security_mmap_file(). I believe that this may be an issue with aufs4 as well, on a RT kernel (with rwsem reader restriction i.e kernels prior to 4.9) . The AB-BA locking order becomes a problem if following happens (where thread A and B are from same process)

      Ok, your version is so old and the commit I wrote
      7aac34b 2014-05-06 aufs: bugfix, stop calling security_mmap_file() again
      is unrelated.

      The commit you mentioned is
      f78f003 2013-03-12 Merge branch 'aufs3.0/10static' into aufs3.0/11proc_map
      right?
      And aufs_mmap() doesn't call security_mmap_file()?
      That is strange. It should have security_mmap_file() call.
      Did you change the aufs source file or apply some patches?

      Since your version is unknown to me, it is difficult for me to suggest
      you a solution. But the approach you wrote first seems bad.

      would it be ok to take mmap_sem lock before taking si_rwsem for all aufs IO calls in case of RT kernel? or will that be causing any other lock races in aufs?

      Rather than that, I'd suggest you to unlock si_rwsem before ->mmap()
      call. It should be better than yours, I guess.

      J. R. Okajima

       
  • Ravindra Khati

    Ravindra Khati - 2018-02-27

    Sorry, my bad, we are certainly based on
    f78f003 2013-03-12 Merge branch 'aufs3.0/10static' into aufs3.0/11proc_map
    and it does have security_mmap_file(), probably I grepped a wrong string earlier.
    We have some trivial patches i.e spin_lock calls changed with seq_spin_lock.

    1) Pardon me for my knowledge on aufs code, I haven't digged in to it much. As you said taking mmap_sem lock before si_rwsem in IO calls would be bad: is that only from a perspective that it's going to delay the mmap calls(i.e waiting for mmap_sem lock) for some threads of a process or is there anything else?

    2) Your suggestion to unlock si_rwsem before mmap(), do you mean abandon the page_fault (which is waiting to acquire the mmap_sem) raised during aufs_read and unlock the si_rwsem so that the other thread doing mmap() could acquire si_rwsem. The rwsem are implemented as recursive mutex in RT kernel, so these can be only unlocked by the thread which has locked it.

    Below is the stack dumps of the threads from same process which are causing the deadlock, I hope that would give a better picture
    ntp_prober D 8031a97c 0 2332 2084 0x00000000
    [<8031a97c>] (__schedule+0x554/0x684) from [<8031ab48>] (schedule+0x9c/0xc4)
    [<8031ab48>] (schedule+0x9c/0xc4) from [<8031bd98>] (__rt_mutex_slowlock+0x84/0xac)
    [<8031bd98>] (__rt_mutex_slowlock+0x84/0xac) from [<8031c1f8>] (rt_mutex_slowlock+0xf4/0x14c)
    [<8031c1f8>] (rt_mutex_slowlock+0xf4/0x14c) from [<8007c46c>] (__rt_down_read.isra.0+0x2c/0x3c)
    [<8007c46c>] (__rt_down_read.isra.0+0x2c/0x3c) from [<8003b140>] (do_page_fault+0x7c/0x2bc)
    [<8003b140>] (do_page_fault+0x7c/0x2bc) from [<80030324>] (do_DataAbort+0x34/0x98)
    [<80030324>] (do_DataAbort+0x34/0x98) from [<80030bd0>] (__dabt_svc+0x70/0xa0)
    Exception stack(0x9cce9d70 to 0x9cce9db8)
    9d60: 2b556000 80805600 00000000 00000000
    9d80: 9cce9e28 0000002c 00000400 00000000 80805600 80471ed0 80805600 0000002c
    9da0: 00000fff 9cce9db8 800995dc 800982b0 200d0013 ffffffff
    [<80030bd0>] (__dabt_svc+0x70/0xa0) from [<800982b0>] (file_read_actor+0x34/0xfc)
    [<800982b0>] (file_read_actor+0x34/0xfc) from [<800995dc>] (generic_file_aio_read+0x3f0/0x6a8)
    [<800995dc>] (generic_file_aio_read+0x3f0/0x6a8) from [<800bf53c>] (do_sync_read+0x98/0xd4)
    [<800bf53c>] (do_sync_read+0x98/0xd4) from [<800bf9d4>] (vfs_read+0xfc/0x188)
    [<800bf9d4>] (vfs_read+0xfc/0x188) from [<801895c4>] (vfsub_read_u+0xc/0x28)
    [<801895c4>] (vfsub_read_u+0xc/0x28) from [<801923e0>] (aufs_read+0x8c/0x104)
    [<801923e0>] (aufs_read+0x8c/0x104) from [<800bf9d4>] (vfs_read+0xfc/0x188)
    [<800bf9d4>] (vfs_read+0xfc/0x188) from [<800bfda4>] (sys_read+0x34/0x68)
    [<800bfda4>] (sys_read+0x34/0x68) from [<800311a0>] (ret_fast_syscall+0x0/0x30)

    ntp_prober D 8031a97c 0 2333 2084 0x00000000
    [<8031a97c>] (__schedule+0x554/0x684) from [<8031ab48>] (schedule+0x9c/0xc4)
    [<8031ab48>] (schedule+0x9c/0xc4) from [<8031bd98>] (__rt_mutex_slowlock+0x84/0xac)
    [<8031bd98>] (__rt_mutex_slowlock+0x84/0xac) from [<8031c1f8>] (rt_mutex_slowlock+0xf4/0x14c)
    [<8031c1f8>] (rt_mutex_slowlock+0xf4/0x14c) from [<8007c46c>] (__rt_down_read.isra.0+0x2c/0x3c)
    [<8007c46c>] (__rt_down_read.isra.0+0x2c/0x3c) from [<80181680>] (si_read_lock+0xa0/0x130)
    [<80181680>] (si_read_lock+0xa0/0x130) from [<80191f50>] (aufs_mmap+0x38/0x298)
    [<80191f50>] (aufs_mmap+0x38/0x298) from [<800b4224>] (mmap_region+0x218/0x418)
    [<800b4224>] (mmap_region+0x218/0x418) from [<800b47d8>] (sys_mmap_pgoff+0x84/0xcc)
    [<800b47d8>] (sys_mmap_pgoff+0x84/0xcc) from [<800311a0>] (ret_fast_syscall+0x0/0x30)

    Thread 2332: Tries to take mmap_sem lock and goes to sleep as it's already locked by thread 2333 in mmap() call

    Thread 2333: Tries to take si_rwsem lock and goes to sleep as it's already locked by thread 2332 in aufs_read().

    Regards,
    Ravindra Khati

     

    Last edit: Ravindra Khati 2018-02-27
    • J. R. Okajima

      J. R. Okajima - 2018-02-27

      "Ravindra Khati":

      1) Pardon me for my knowledge on aufs code, I haven't digged in to it much. As you said taking mmap_sem lock before si_rwsem in IO calls would be bad: is that only from a perspective that it's going to delay the mmap calls(i.e waiting for mmap_sem lock) for some threads of a process or is there anything else?

      si_rwsem protects au_sbinfo members. The typical case is a branch
      management such as to add/del/mod branches. I am afraid that you don't
      use such feature.

      2) Your suggestion to unlock si_rwsem before mmap(), do you mean abandon the page_fault (which is waiting to acquire the mmap_sem) raised during aufs_read and unlock the si_rwsem so that the other thread doing mmap() could acquire si_rwsem. The rwsem are implemented as recursive mutex in RT kernel, so these can be only unlocked by the thread which has locked it.

      Do you mean read(2) systemcall issues mmap(2) internally?
      I am not sure there such path exists... But I thought page-fault will
      call address_space_operations instead of file_operations.
      And I didn't mean disturbing page_fault. Just move forward/upward
      si_read_unlock(sb) call before ->mmap() call in aufs_mmap(). As long as
      other tasks don't change the member in au_sbinfo, it should be safe.

      J. R. Okajima

       
  • Ravindra Khati

    Ravindra Khati - 2018-02-28

    si_rwsem protects au_sbinfo members. The typical case is a branch
    management such as to add/del/mod branches. I am afraid that you don't
    use such feature.

    We do have branches added which we do during boot up, but yes we don't delete or modify later and do not plan to do so.

    Do you mean read(2) systemcall issues mmap(2) internally?
    I am not sure there such path exists... But I thought page-fault will
    call address_space_operations instead of file_operations.
    And I didn't mean disturbing page_fault.

    No, I didn't mean that. What I am saying is, there are two separate threads in user space process of which, one is doing a read() system call and another is doing mmap2() system call and there is a race between these two threads for locks.
    Here is the flow in kernel code for these two as they are runnig parallely.
    1) Thread A: Doing read() sys call to read content of some file, here a mmapped buffer is passed which is causing a page fault: read()-->aufs_read()-->si_read_lock() lock acquired-->vfsub_read_u-->generic_file_aio_read()->page_fault-->trying to lock mmap_sem (blocked) as it's locked by "thread B"
    2) Thread B: is doing mmap2() syscall to get some buffer
    mmap2()->taken mmap_sem write lock(mm/mmap.c)-->aufs_mmap()-->si_read_lock()--> (blocked) as si_read_lock is acquired by thread A.

    Just move forward/upward
    si_read_unlock(sb) call before ->mmap() call in aufs_mmap(). As long as
    other tasks don't change the member in au_sbinfo, it should be safe.

    I suppose, you are reffering to h_file->f_op->mmap() call in aufs_mmap()? this isn't causing any issue here. The thread doing mmap2() sys call is blocked at the si_read_lock() in aufs_mmap, way before this internal ->mmap() call in aufs_mmap.

    Regards,
    Ravindra Khati

     
    • J. R. Okajima

      J. R. Okajima - 2018-02-28

      "Ravindra Khati":

      1) Thread A: Doing read() sys call to read content of some file, here a mmapped buffer is passed which is causing a page fault: read()-->aufs_read()-->si_read_lock() lock acquired-->vfsub_read_u-->generic_file_aio_read()->page_fault-->trying to lock mmap_sem (blocked) as it's locked by "thread B"
      2) Thread B: is doing mmap2() syscall to get some buffer
      mmap2()->taken mmap_sem write lock(mm/mmap.c)-->aufs_mmap()-->si_read_lock()--> (blocked) as si_read_lock is acquired by thread A.

      "read-lock" for rwsem (aka. "shared-lock"), it is possible for TWO
      thread to acquire the single rwsem, isn't it? Or your RT kernel doesn't
      allow such shared-lock?

      J. R. Okajima

       
  • Ravindra Khati

    Ravindra Khati - 2018-02-28

    "read-lock" for rwsem (aka. "shared-lock"), it is possible for TWO
    thread to acquire the single rwsem, isn't it? Or your RT kernel doesn't
    allow such shared-lock?

    That's the whole point, rwsem is unfortunately implemented as a recursive mutex in standard RT kernel. One thread can lock it multiple times but no two threads can have the lock simultaneously (not even read-lock :( ). This is done deliberatley in RT kernel so that the writers do not starve on readers. Since it's a typical issue of RT kernel + aufs, I believe this may be true for aufs4 as well, as I can see that locking sequence is same. In RT kernel 4.9, I can see this restriction on rwsem has been lifted off.
    1) So either I need to backport that patch to our old RT kernel, I will have to see how feasible is that and will have larger impact on our system.
    2) or have the mmap_sem locked in aufs_read/aufs_write before taking si_read_lock (at least in case of multithreaded programs). That certainly would make the locking order similar to mmap2() case thus avoiding deadlock. But then if I understood your earlier response correctly, this is not good from branch maintainenance pov? isn't it?

     
    • J. R. Okajima

      J. R. Okajima - 2018-02-28

      "Ravindra Khati":

      That's the whole point, rwsem is unfortunately implemented as a recursive mutex in standard RT kernel. One thread can lock it multiple times but no two threads can have the lock simultaneously (not even read-lock :( ). This is done deliberatley in RT kernel so that the writers do not starve on readers. Since it's a typical issue of RT kernel + aufs, I believe this may be true for aufs4 as well, as I can see that locking sequence is same. In RT kernel 4.9, I can see this restriction on rwsem has been lifted off.

      Ok, now I think I got your point.

      2) or have the mmap_sem locked in aufs_read/aufs_write before taking si_read_lock (at least in case of multithreaded programs). That certainly would make the locking order similar to mmap2() case thus avoiding deadlock. But then if I understood your earlier response correctly, this is not good from branch maintainenance pov? isn't it?

      I didn't know such rwsem behaviour in RT kernel, and I thought mmap_sem
      would cause a problem such like deadlock. As you wrote, if a task can
      acquire mmap_sem multiple times, then it will be safe and I'd agree with
      you. mmap_sem before si_rwsem is worth to try.

      And if you know the detail about 4.9 RT kernel, would you tell me how to
      reproduce the problem on my side? I have vanilla mainline 4.9 and latest
      aufs4.9 (off course). And what do I need?

      J. R. Okajima

       
  • Ravindra Khati

    Ravindra Khati - 2018-03-01

    As you wrote, if a task can
    acquire mmap_sem multiple times, then it will be safe and I'd agree with
    you. mmap_sem before si_rwsem is worth to try.

    Yes, I have tried this for the race between aufs_read and mmap, this seems to fix the issue.
    but I think I would have to put this fix in all the calls where si_read_lock is being called and there are chances of page fault being raised raised in that flow.

    And if you know the detail about 4.9 RT kernel, would you tell me how to
    reproduce the problem on my side? I have vanilla mainline 4.9 and latest
    aufs4.9 (off course). And what do I need?

    here is the 4.9-20-rt15 kernel tar, this should have the restriction on rwsem lock
    https://git.kernel.org/pub/scm/linux/kernel/git/rt/linux-stable-rt.git/commit/?h=v4.9-rt&id=702775a5e8163bb5199b917081f33a9642e72966

    I got this issue by calling a getaddrinfo_a (gnu c lib, dns query function), which supposedly is creating multiple threads and causing this race to occur. But this doesn't happen every time, I just run a loop in shell to run the binary multiple times. I am also trying to reproduce this with a simple multithreaded program, doing mmap and read calls in each threads but so far I am not able to reproduce it this way.

    Ravindra Khati

     

Log in to post a comment.