OpenSAF / Tickets / #2094 Standby controller goes for reboot on stopping openSaf with STONITH enabled cluster

Hans Nordebäck - 2016-10-06

This is the same behaviour as running without stontih or PLM. Without stonith opensaf tries to reboot the standby controller at opensafd stop, but needs either PLM or stonith to succeed. Perhaps it is needed to stop opensaf and not trigger remote fencing? Is this an upgrade case? Perhaps we should create an enhancement ticket for this?

If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Mathi Naickan - 2016-10-06

This seems to be a case of differentiating a hung node versus a node on which the middleware is stopped.

Is there any standard means to detect a "hung" node?
IF there is such a mechanism to detect a hung node, then
Upon receiving "NODE_DOWN" i.e. below event
"Oct 5 13:01:24 SC-1 osaffmd[5526]: NO Node Down event for node id 2020f:"
FM could use (say a libvirt command) command to detect if the node is hung or running healthy. If running healthy then a reboot using stonith could be avoided.

OpenSAF did support the usecase of "/etc/init.d/opensafd stop without OS reboot". Should we continue to support that?

If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Anders Widell - 2016-10-06

I think the procedure for stopping OpenSAF in a controlled way is to first lock the node using CLM. The CLM lock admin operation will remove the node from cluster membership. The it should be safe to stop OpenSAF on that node without getting fenced - i.e. we should not fence a node that we lost contact with if the node was not a member of the cluster.

If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Chani Srivastava - 2016-10-13

Is Stonith applicable only for controllers? As no reboot observed while stopping opensaf on Payload.

If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Hans Nordebäck - 2016-10-13

Split brain may only happen between the system controllers.

If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Chani Srivastava - 2016-11-02

Can you provide the documentation on how to stop opensaf in a controlled manner so that I can close the ticket.

If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Hans Nordebäck - 2016-11-02

Ticket [#2160] will add support to differentiate between a hung versus a stopped node, no additional documentation will be needed.

Related

Tickets: ~~#2160~~

If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Srikanth R - 2016-11-07

There are two scenarios where "opensafd stop" is invoked on any opensaf controller.

SCENARIO-1) Where /etc/init.d/opensafd script is invoked manually on command prompt when the system is running and up.
SCENARIO-2) Software on a controller ( other than opensafd) invoked "reboot" for which opensafd stop is invoked in run level 3 or higher.

With the patch submitted for #2160,

a)node shall go for reboot in scenario-1, if administrator doesn't invoke clm admin operation. This is fine.

b) For scenario-2, all run level services shall not be stopped gracefully as the node shall be rebooted abruptly after opensafd stop as admin did not invoke clm admin operation. So, opensafd as a HA software shall not support graceful reboot on standby controller with the #2160 fix ?

If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Hans Nordebäck - 2016-11-08

in a) doing /etc/init.d/opensafd stop doesn't reboot the node, but stops opensaf on that node and saClmNodeIsMember is set to false. The active controller will then not perform remote fencing of that node.
in b) "graceful" reboot after opensafd stop, should work fine without any involvemnet of the remote fencing functionality

If you would like to refer to this comment somewhere else in this project, copy and paste the following link:
- Srikanth R - 2016-11-08
  
  For the scenario-2,
  -> Management software e.g. SWN other than opensafd issued reboot on standby controller. From opensaf perspecitve , the standby controller might be healthy member of a cluster. But from the SWN perspecitve, node needs to be repaired and reboot is invoked.
  
  -> When reboot command is invoked by SWN, all services in configured runlevel shall be stopped in the order.
  
  -> Once the opensafd stop script is invoked on standby controller, active controller detects that the standby controller is in healthy state and remote fencing shall be done.
  
  -> As part of remote fencing, the node shall be hard rebooted, which doesn't give chance for other services in runlevel to be stopped gracefully.
  
  -> If the SWN has a database service ( e.g. drbd) which is to be stopped after opensafd stop, the database service stop script shall not be invoked as remote fencing is done. This may result in bad state for the other management software e.g. SWN.
  
  Suggestion :
  
  1) Either opensaf shall document that admin needs to perform clm admin lock of standby controller before repairing. OR
  2) FM should detect the difference between opensafd stop and hung opensaf processes. As part of opensafd stop, peer fmd on standby contoller can update fmd on active controller that opensafd on standby is going gracefully.
  
  If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Hans Nordebäck - 2016-11-11

Agree, "Suggestion: 1" document that admin needs to perform clm admin lock of standby is a good suggestion. The node will then not be a member of the cluster and not affected by remote fencing

If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Anders Widell - 2017-02-28

Milestone: 5.2.FC --> 5.2.RC1
If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Hans Nordebäck - 2017-03-01

I suggest to close this ticket as a duplicate of ticket #2160

If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Chani Srivastava - 2017-03-01

status: unassigned --> duplicate
If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Chani Srivastava - 2017-03-01

Closing as duplicate of #2160

If you would like to refer to this comment somewhere else in this project, copy and paste the following link:

Standby controller goes for reboot on stopping openSaf with STONITH enabled cluster

Milestone

Searches

Help

#2094 Standby controller goes for reboot on stopping openSaf with STONITH enabled cluster

Related

Discussion

Related