Get the App
SLTechnology News&Howtos  ›  Database  › 

Evicting instance 2 from cluster caused by private Network instability

Shulou Source: shulou.com Published: 2022-06-01 17:41:48 10月01日 Update

Environment: dual-node RAC, oracle 11.2.3

Customer phone RAC instance 2 is abnormal. Check the log on the spot:

Example 2:

Fri Aug 25 09:45:16 2017

Received an instance abort message from instance 1Received an instance abort message from instance 1

Please check instance 1 alert and LMON trace files for detail.Please check instance 1 alert and LMON trace files for detail.

LMS0 (ospid: 24510820): terminating the instance due to error

Fri Aug 25 09:45:16 2017

System state dump requested by (instance=2, osid=24510820 (LMS0)), summary= [abnormal instance termination].

System State dumped to trace file / oracle/11.2.0/diag/rdbms/ins/ins2/trace/ins2_diag_21561818.trc

Instance terminated by LMS0, pid = 24510820

Example 1

Fri Aug 25 09:44:25 2017

IPC Send timeout detected. Sender: ospid 35783054 [oracle@db1 (LMS1)]

Receiver: inst 2 binc 2073329022 ospid 24183072

IPC Send timeout to 2.2 inc 28 for msg type 65518 from opid 14

Fri Aug 25 09:44:27 2017

Communications reconfiguration: instance_number 2

Fri Aug 25 09:45:16 2017

Detected an inconsistent instance membership by instance 1

Evicting instance 2 from cluster

Waiting for instances to leave: 2

Fri Aug 25 09:45:16 2017

Dumping diagnostic data in directory= [CDMP _ 20170825094516], requested by (instance=2, osid=24510820 (LMS0)), summary= [abnormal instance termination].

Reconfiguration started (old inc 28, new inc 32)

List of instances:

1 (myinst: 1)

View / oracle/11.2.0/diag/rdbms/gjj/ins2/trace/ins2_diag_21561818.trc

* 2017-08-25 14 14 24 purl 35.900

I'm the voting node

Group reconfiguration cleanup

Confirm- > incar_num 22, rcfgctx- > prop_incar 0

Send my bitmap to master 0

Kjzgmappropose: incar 0, newmap-

3000000000000000000000000000000000000000000000000000000000000000

Kjzgmappropose: rc from psnd: 30

Kjzdattdlm: Can not attach to DLM (LMON up= [TRUE], DB mounted= [FALSE]).

Kjzdattdlm: Can not attach to DLM (LMON up= [TRUE], DB mounted= [FALSE]).

It is suspected that there is a problem with the heartbeat network (this set of RAC has been expelled several times before, but the instance is automatically activated. Instance 2 cannot be started after this instance is expelled. The parameters have been modified to solve the problem of the previous instance being expelled. In this case, it is not a parameter setting problem).

Test heartbeat network, connectivity and transmission rate are not problems, the subsequent plan to further improve heartbeat network availability through haip, in the process of adding haip found that when the server and the server and switch newly added network after the packet loss, packet loss rate of 50%, determine the heartbeat network stability problems, based on this to remove the newly added heartbeat line, replace the original heartbeat line, restart the expelled instance 2 The instance is normal.

Finally, it is judged that the fault is caused by a two-core short circuit in the original heartbeat line RJ45 head.

Tags: Instance problem network parameter situation switch availability customer failure data log server environment phone stability node process rate one transmission Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno MariaDB vpn Huawei Shulou Tech Info Xiaomi