Menu ▾ ▴

#809 AMF: ensure IMM is updated before notification is sent

future
unassigned
nobody
None
enhancement
amf
d
4.4.RC2
major
2014-04-10
2014-03-11
Shu Wang
No

We observed a delay of between the notification receipt of SA_NTF_TYPE_STATE_CHANGE and IMM update. Our code watches for SI state change notifications and needs to query the IMM to get the assigned SU. When the query is performed the data does not exist in the IMM.

Looking at the OpenSAF code, when unlock of the SU is performed , the OpenSAF code handled this is avd_su_admin_state_set() in su.cc
void avd_su_admin_state_set(AVD_SU *su, SaAmfAdminStateT admin_state)
{
... ...
avd_saImmOiRtObjectUpdate(&su->name,
const_cast<saimmattrnamet>("saAmfSUAdminState"), SA_IMM_ATTR_SAUINT32T, &su->saAmfSUAdminState);
m_AVSV_SEND_CKPT_UPDT_ASYNC_UPDT(avd_cb, su, AVSV_CKPT_SU_ADMIN_STATE);
avd_send_admin_state_chg_ntf(&su->name, SA_AMF_NTFID_SU_ADMIN_STATE, old_state, su->saAmfSUAdminState);
}</saimmattrnamet>

avd_saImmOiRtObjectUpdate() does not process the job immediately, the job is added to a queue. Then avd_send_admin_state_chg_ntf() sends the notification.

Dequeue happens in main_loop() in main.cc:
static void main_loop(void)
{
... ...
while(1) {
int pollretval = poll(fds, nfds, polltmo);
......
if (cb->immOiHandle && fds[FD_IMM].revents & POLLIN) {
TRACE("IMM event rec");
error = saImmOiDispatch(cb->immOiHandle, SA_DISPATCH_ALL);
... ...
// If there is no error
/ commit async updated possibly sent in the callback /
m_AVSV_SEND_CKPT_UPDT_SYNC(cb, NCS_MBCSV_ACT_UPDATE, 0);
/ flush messages possibly queued in the callback /
avd_d2n_msg_dequeue(cb);
}
// submit some jobs (if any)
polltmo = retval_to_polltmo(Fifo::execute(cb->immOiHandle));
}
}

From trace, we found the saImmOiDispatch() call took some time and it delayed the dequeue.
Is it possible to have a separate thread that is processing the dequeue jobs so that the imm is updating more quickly.

Trace is attached in attachment.

We use Oracle linux version: 2.6.39-400.17.1.el6uek.x86_64

Unlock a SU can reproduce the problem.

1 Attachments

Related

Tickets: #809

Discussion

  • Shu Wang

    Shu Wang - 2014-03-11
    • Description has changed:

    Diff:

    --- old
    +++ new
    @@ -35,8 +35,10 @@
     }
    
     From trace, we found the saImmOiDispatch() call took some time and it delayed the dequeue.
    +Is it possible to have a separate thread that is processing the dequeue jobs so that the imm is updating more quickly.
    +
     Trace is attached in attachment.
    
     We use Oracle linux version: 2.6.39-400.17.1.el6uek.x86_64
    +
     Unlock a SU can reproduce the problem.
    -
    
     
  • Anders Bjornerstedt

    Using NTF notifications as an alternative to the "applier" interface will
    work in the context of NTF. The notification should have all the information
    about the change, so there should be no need to perform a redundant and
    UNSAFE read towards the IMM.

    The IMM uses a broadcast protocol called fevs to provide synhcrony across
    the cluster. Other services such as NTF do not use fevs.

    Thus you can NOT send a message from say an OI in an apply callback
    on one processor, to a receiver process at another processor using some
    other arbitrary messaging mechanism and expect the receiver processor
    to be in sync with the imm state at that other node.

    So either use NTF and only NTF. Or use the imm and only the imm.
    Instead of using NTF you could use the imm applier interface.

    There is a way for an application to send imm-synchronous (fevs-synchronous)
    messages to a receiver process at another processor: By using an
    admin-operation, which goes over imm-fevs. So an OI or applier could
    send an admin-operation from inside an apply callback to a receiver
    process at another processor and be sure that the CCB had been applied
    at the other processor by the time they received the admiin-op
    (this is backwards synchrony). But doing so will only guarantee
    that the admin-op arrives after the ccb has also been applied there.
    The problem is that in theory you could also have a subsequent CCB arriving
    before the admin op that overwrites the same objects written by the first
    CCB.

    So the only mechanism to be exactly informed about the change of a CCB is
    to use the applier interface (or the NTF IMM notifications exclusively).

     
    • Shu Wang

      Shu Wang - 2014-03-11

      Thanks for the reply. We would like to use the NTF but the NTF does not have enough information. We have processes that watch when other SUs are assigned a SI and when the SI is removed. When we unlock a process (SI is assigned), the following notification is the only one received:

      eventType = SA_NTF_OBJECT_STATE_CHANGE
      notificationObject = "safSi=amfRaterSI1.4,safApp=olcApp"
      notifyingObject = "safApp=safAmfService"
      notificationClassId = SA_NTF_VENDOR_ID_SAF.SA_SVC_AMF.111 (0x6f)
      additionalText = "The Assignment state of SI safSi=amfRaterSI1.4,safApp=olcApp changed"
      sourceIndicator = SA_NTF_OBJECT_OPERATION
      State ID = SA_AMF_ASSIGNMENT_STATE
      New State: SA_AMF_ASSIGNMENT_FULLY_ASSIGNED

      The NTF contains no information about what SU that the assignment is related to. The only way to find that is to read the IMM. If a different NTF was received related to the SU we would be OK.
      On the lock (SI removed), we do receive a HA State change notification that does contain the SU so in this case the NTF does contain enough information for us (contains the SI and SU):

      eventType = SA_NTF_OBJECT_STATE_CHANGE
      notificationObject = "safSu=amfRaterSU1.2,safSg=amfRaterSG1,safApp=olcApp"
      notifyingObject = "safApp=safAmfService"
      notificationClassId = SA_NTF_VENDOR_ID_SAF.SA_SVC_AMF.110 (0x6e)
      additionalText = "The HA state of SI safSi=amfRaterSI1.4,safApp=olcApp assigned to SU safSu=amfRaterSU1.2,safSg=amfRaterSG1,safApp=olcApp changed"

      • additionalInfo: 0 -
        infoId = 2
        infoType = 10
        infoValue = "safSi=amfRaterSI1.4,safApp=olcApp"
        sourceIndicator = SA_NTF_OBJECT_OPERATION
        State ID = SA_AMF_HA_STATE
        New State: SA_AMF_HA_QUIESCED

      Should we change the defect to state that a HA state change notification should be produced when unlocking a SU and it is assigned a SI?

      Shu Wang | Senior Analyst | +1(407)708-5117 or x3917| www.NetCracker.com
      Proven Partner to Communications Service Providers

      From: Anders Bjornerstedt [mailto:andersbj@users.sf.net]
      Sent: Tuesday, March 11, 2014 9:32 AM
      To: [opensaf:tickets]
      Subject: [opensaf:tickets] #809 Delay between notification and IMM update when admin state changes

      Using NTF notifications as an alternative to the "applier" interface will
      work in the context of NTF. The notification should have all the information
      about the change, so there should be no need to perform a redundant and
      UNSAFE read towards the IMM.

      The IMM uses a broadcast protocol called fevs to provide synhcrony across
      the cluster. Other services such as NTF do not use fevs.

      Thus you can NOT send a message from say an OI in an apply callback
      on one processor, to a receiver process at another processor using some
      other arbitrary messaging mechanism and expect the receiver processor
      to be in sync with the imm state at that other node.

      So either use NTF and only NTF. Or use the imm and only the imm.
      Instead of using NTF you could use the imm applier interface.

      There is a way for an application to send imm-synchronous (fevs-synchronous)
      messages to a receiver process at another processor: By using an
      admin-operation, which goes over imm-fevs. So an OI or applier could
      send an admin-operation from inside an apply callback to a receiver
      process at another processor and be sure that the CCB had been applied
      at the other processor by the time they received the admiin-op
      (this is backwards synchrony). But doing so will only guarantee
      that the admin-op arrives after the ccb has also been applied there.
      The problem is that in theory you could also have a subsequent CCB arriving
      before the admin op that overwrites the same objects written by the first
      CCB.

      So the only mechanism to be exactly informed about the change of a CCB is
      to use the applier interface (or the NTF IMM notifications exclusively).


      [tickets:#809]http://sourceforge.net/p/opensaf/tickets/809/ Delay between notification and IMM update when admin state changes

      Status: unassigned
      Milestone: future
      Created: Tue Mar 11, 2014 01:12 PM UTC by Shu Wang
      Last Updated: Tue Mar 11, 2014 01:17 PM UTC
      Owner: nobody

      We observed a delay of between the notification receipt of SA_NTF_TYPE_STATE_CHANGE and IMM update. Our code watches for SI state change notifications and needs to query the IMM to get the assigned SU. When the query is performed the data does not exist in the IMM.

      Looking at the OpenSAF code, when unlock of the SU is performed , the OpenSAF code handled this is avd_su_admin_state_set() in su.cc
      void avd_su_admin_state_set(AVD_SU *su, SaAmfAdminStateT admin_state)
      {
      ... ...
      avd_saImmOiRtObjectUpdate(&su->name,
      const_cast("saAmfSUAdminState"), SA_IMM_ATTR_SAUINT32T, &su->saAmfSUAdminState);
      m_AVSV_SEND_CKPT_UPDT_ASYNC_UPDT(avd_cb, su, AVSV_CKPT_SU_ADMIN_STATE);
      avd_send_admin_state_chg_ntf(&su->name, SA_AMF_NTFID_SU_ADMIN_STATE, old_state, su->saAmfSUAdminState);
      }

      avd_saImmOiRtObjectUpdate() does not process the job immediately, the job is added to a queue. Then avd_send_admin_state_chg_ntf() sends the notification.

      Dequeue happens in main_loop() in main.cc:
      static void main_loop(void)
      {
      ... ...
      while(1) {
      int pollretval = poll(fds, nfds, polltmo);
      ......
      if (cb->immOiHandle && fds[FD_IMM].revents & POLLIN) {
      TRACE("IMM event rec");
      error = saImmOiDispatch(cb->immOiHandle, SA_DISPATCH_ALL);
      ... ...
      // If there is no error
      / commit async updated possibly sent in the callback /
      m_AVSV_SEND_CKPT_UPDT_SYNC(cb, NCS_MBCSV_ACT_UPDATE, 0);
      / flush messages possibly queued in the callback /
      avd_d2n_msg_dequeue(cb);
      }
      // submit some jobs (if any)
      polltmo = retval_to_polltmo(Fifo::execute(cb->immOiHandle));
      }
      }

      From trace, we found the saImmOiDispatch() call took some time and it delayed the dequeue.
      Is it possible to have a separate thread that is processing the dequeue jobs so that the imm is updating more quickly.

      Trace is attached in attachment.

      We use Oracle linux version: 2.6.39-400.17.1.el6uek.x86_64

      Unlock a SU can reproduce the problem.


      Sent from sourceforge.net because you indicated interest in https://sourceforge.net/p/opensaf/tickets/809/

      To unsubscribe from further messages, please visit https://sourceforge.net/auth/subscriptions/


      The information transmitted herein is intended only for the person or entity to which it is addressed and may contain confidential, proprietary and/or privileged material. Any review, retransmission, dissemination or other use of, or taking of any action in reliance upon, this information by persons or entities other than the intended recipient is prohibited. If you received this in error, please contact the sender and delete the material from any computer.

       

      Related

      Tickets: #809

  • Anders Bjornerstedt

    • status: unassigned --> invalid
     
  • Anders Bjornerstedt

    I think you need to re-analyze what it is exactly that you are trying to
    accomplish. In general it is not possible to capture a distributed system
    with many processes on many processors as being in "one state".

    Most systems tend to have a stable configuration during long periods.
    But then there may be a flurry of several CCBs changing imm config data,
    which then ripples through the components of one or more services.
    Each component may be slower or faster in processing the changes.
    One can not expect everyone to wait on everyone for ack on completion of
    every step of a change.

    Instead try to define what exactly it is you need to capture. Derived from
    what you need to accomplish. Not more and not less. Once you know what you
    are trying to capture, then you can try to solve the problem of how to do
    it.

    One thing is to capture when a certain kind of change starts. Another thing
    is to define and capture when a change is completed. Not forgetting here
    that you may hae several changes in tight sequence.

     
    • Shu Wang

      Shu Wang - 2014-03-11

      Thanks for the response. To give you some background we are converting from using SAFfire to OpenSAF. Our cluster/AIS logic works with SAFfire - it has been working with SAFfire for years. We are struggling to get our cluster working with OpenSAF. The notification related logic is one area that we are having problems with. We are able to watch for notifications related to the assignment/removal of a SI to a SU.
      We have a distributor process that sends events to a receiver. If the receiver is not in an assigned state (no SI assigned to it), the distributor should not send events to it. On initial startup, we read the IMM to understand the state of our receivers/SUs. We then rely on notifications to understand when the SI assignments/removals occur for those objects. The distributor needs to understand when our receivers are able to do work and when they are not (in real-time, we need to process thousands of events per second). We don't want to send an event to a receiver that can't process it. We understand that many changes can occur in a short period of time and have coded for that. Our struggle is to get information from OpenSAF to understand when a receiver is assigned. There is a notification on the unassignment.

      Shu Wang | Senior Analyst | +1(407)708-5117 or x3917| www.NetCracker.com
      Proven Partner to Communications Service Providers

      From: Anders Bjornerstedt [mailto:andersbj@users.sf.net]
      Sent: Tuesday, March 11, 2014 11:03 AM
      To: [opensaf:tickets]
      Subject: [opensaf:tickets] #809 Delay between notification and IMM update when admin state changes

      I think you need to re-analyze what it is exactly that you are trying to
      accomplish. In general it is not possible to capture a distributed system
      with many processes on many processors as being in "one state".

      Most systems tend to have a stable configuration during long periods.
      But then there may be a flurry of several CCBs changing imm config data,
      which then ripples through the components of one or more services.
      Each component may be slower or faster in processing the changes.
      One can not expect everyone to wait on everyone for ack on completion of
      every step of a change.

      Instead try to define what exactly it is you need to capture. Derived from
      what you need to accomplish. Not more and not less. Once you know what you
      are trying to capture, then you can try to solve the problem of how to do
      it.

      One thing is to capture when a certain kind of change starts. Another thing
      is to define and capture when a change is completed. Not forgetting here
      that you may hae several changes in tight sequence.


      [tickets:#809]http://sourceforge.net/p/opensaf/tickets/809/ Delay between notification and IMM update when admin state changes

      Status: invalid
      Milestone: future
      Created: Tue Mar 11, 2014 01:12 PM UTC by Shu Wang
      Last Updated: Tue Mar 11, 2014 01:33 PM UTC
      Owner: nobody

      We observed a delay of between the notification receipt of SA_NTF_TYPE_STATE_CHANGE and IMM update. Our code watches for SI state change notifications and needs to query the IMM to get the assigned SU. When the query is performed the data does not exist in the IMM.

      Looking at the OpenSAF code, when unlock of the SU is performed , the OpenSAF code handled this is avd_su_admin_state_set() in su.cc
      void avd_su_admin_state_set(AVD_SU *su, SaAmfAdminStateT admin_state)
      {
      ... ...
      avd_saImmOiRtObjectUpdate(&su->name,
      const_cast("saAmfSUAdminState"), SA_IMM_ATTR_SAUINT32T, &su->saAmfSUAdminState);
      m_AVSV_SEND_CKPT_UPDT_ASYNC_UPDT(avd_cb, su, AVSV_CKPT_SU_ADMIN_STATE);
      avd_send_admin_state_chg_ntf(&su->name, SA_AMF_NTFID_SU_ADMIN_STATE, old_state, su->saAmfSUAdminState);
      }

      avd_saImmOiRtObjectUpdate() does not process the job immediately, the job is added to a queue. Then avd_send_admin_state_chg_ntf() sends the notification.

      Dequeue happens in main_loop() in main.cc:
      static void main_loop(void)
      {
      ... ...
      while(1) {
      int pollretval = poll(fds, nfds, polltmo);
      ......
      if (cb->immOiHandle && fds[FD_IMM].revents & POLLIN) {
      TRACE("IMM event rec");
      error = saImmOiDispatch(cb->immOiHandle, SA_DISPATCH_ALL);
      ... ...
      // If there is no error
      / commit async updated possibly sent in the callback /
      m_AVSV_SEND_CKPT_UPDT_SYNC(cb, NCS_MBCSV_ACT_UPDATE, 0);
      / flush messages possibly queued in the callback /
      avd_d2n_msg_dequeue(cb);
      }
      // submit some jobs (if any)
      polltmo = retval_to_polltmo(Fifo::execute(cb->immOiHandle));
      }
      }

      From trace, we found the saImmOiDispatch() call took some time and it delayed the dequeue.
      Is it possible to have a separate thread that is processing the dequeue jobs so that the imm is updating more quickly.

      Trace is attached in attachment.

      We use Oracle linux version: 2.6.39-400.17.1.el6uek.x86_64

      Unlock a SU can reproduce the problem.


      Sent from sourceforge.net because you indicated interest in https://sourceforge.net/p/opensaf/tickets/809/

      To unsubscribe from further messages, please visit https://sourceforge.net/auth/subscriptions/


      The information transmitted herein is intended only for the person or entity to which it is addressed and may contain confidential, proprietary and/or privileged material. Any review, retransmission, dissemination or other use of, or taking of any action in reliance upon, this information by persons or entities other than the intended recipient is prohibited. If you received this in error, please contact the sender and delete the material from any computer.

       

      Related

      Tickets: #809

  • Hans Feldt

    Hans Feldt - 2014-03-11

    Maybe the problem can be solved with AMF protection groups?

    There is a small improvement that can be done in AMF but it is more related to IMM updates vs admin operation response. As implemented now there can be a time gap between those to. This is touched upon in the description of this ticket.

    I think there already is a ticket for this enhancement, will search. I also would like this implemented because it would simplify testing of AMF.

     
  • Anders Bjornerstedt

    In this kind of situation, where you know that something is supposed to
    change (in the AMF), but dont have a direct trigger for that change,
    only an indirect trigger that is relatively close (hopefully) in real-time,
    maybe there is some polling (try-again) solution. Perhaps there is some
    AMF runtime attribute that reflects what you are waiting for. Or an admin
    operation ? Or the existence of an object? An SA_AIS_ERR_NOT_EXIST can in
    some cases be handled as an SA_AIS_ERR_TRY_AGAIN if you know the thing
    should exists "soon".

     
  • Anders Bjornerstedt

    If you want a trigger for when a configuration change has been applied
    at a particular processor, then the IMM applier interface is what you
    want to use. The applier interface is a specialization of the OI interface.
    See osaf/services/saf/immsv/README or the OpenSAF_IMMSv_PR doc.

    But a config change typically also means that the service that the
    config data manages, also needs to change state and behavior to reflect
    the config change. If you want a trigger for that change in the service
    (AMF in this case), then you want a notification from the AMF.
    The only alternative would be to poll some AMF runtime attribute,
    or poll using some AMF admin-op that can provide this information.

     
  • Mathi Naickan

    Mathi Naickan - 2014-03-17

    "During the unlock(assigned)....The NTF contains no information about what SU that the assignment is related to. "
    That, I think is still a valid topic of discussion.

    May be there is already an existing enhancement ticket around this, as mentioned by Hans...

     
  • Anders Bjornerstedt

    Note: this ticket (#809) is closed as invalid, but the creator of this ticket
    has created a new ticket (#813) with a reformulated and valid problem
    description.

    Any further comments on this issue is better to put in #813

    https://sourceforge.net/p/opensaf/tickets/813/

     
    • Hans Feldt

      Hans Feldt - 2014-03-18

      Well you closed it as invalid despite it being an AMF ticket...

      This is a reasonable improvement in AMF that I would like to do, the question is if it is a bug or not. Since this is undocumented AMF behavior it could be seen as a defect. I think SMF could benefit from this. But for sure AMF testing would be simplified. Since I haven't found any other ticket I will reopen this one.

       
  • Hans Feldt

    Hans Feldt - 2014-03-18
    • status: invalid --> unassigned
     
  • Anders Bjornerstedt

    I closed it because the problem formulation was not correct.
    They received a notification originating from the IMM applier
    mechanism and presumed that this notification should be in sync
    with a change of state inside the AMF. That was, is, and will always
    be an incorrect assumption.

    If you are planning to do some "reasonable improvement" to the AMF then
    I would call that an enhancement and what that enhancement would be is not
    defined in this ticket.

    You also have ticket #813 currently also defined as a defect.
    Is that ticket addressing a different problem?

     
  • Hans Feldt

    Hans Feldt - 2014-04-10
    • summary: Delay between notification and IMM update when admin state changes --> AMF: ensure IMM is updated before notification is sent
    • Type: defect --> enhancement
     
  • Hans Feldt

    Hans Feldt - 2014-04-10

    Yes the delay is there by design. There are no guarantees that IMM is updated before the notification is received. Maybe if the notification was sent using the same job queue as for IMM updates. A similar change is done for admin op responses in https://sourceforge.net/p/opensaf/tickets/817/

    Changing this to an enhancement.

    There is a connection to a defect where AMF is loosing notification during controller failover. Not sure if that is solved by such change.

     

Log in to post a comment.