How to get rid of broken nodes and add new nodes by etcd
Etcd has been running for a period of time, recently found that etcd alarm, found that an etcd container is damaged, but etcd does not automatically remove the node; at the same time, K8s added a node, but the node has joined the cluster and needs to be re-added.
[root@A01-R04-I69-12222] # etcdctl-endpoints= "http://etcd-xxxl:4001" cluster-health
Failed to check the health of member ec292d985b723e4 on http://10.187.27.196:4001: Get http://10.187.27.196:4001/health: dial tcp 10.187.27.196:4001: getsockopt: connection timed out
Member ec292d985b723e4 is unreachable: [http://10.187.27.196:4001] are all unreachable
Member 2727dc2f519c6794 is healthy: got healthy result from http://10.185.243.35:4001
Member 4e56d8229082190f is healthy: got healthy result from http://10.187.24.132:4001
Member 5780625be722ce57 is healthy: got healthy result from http://10.187.24.134:4001
Member ec2aa2bbe0c891b7 is healthy: got healthy result from http://10.187.27.200:4001
10.187.27.196 was found to be damaged. Record the id ec292d985b723e4 of the node.
Execute the delete command:
Etcdctl-endpoints= "http://etcd-xxxl:4001" member remove ec292d985b723e4
Removed member ec292d985b723e4 from cluster
Check again, the damaged node has been removed.
Etcdctl-endpoints= "http://etcd-xxxxxxx:4001" cluster-health
Member 2727dc2f519c6794 is healthy: got healthy result from http://10.185.243.35:4001
Member 4e56d8229082190f is healthy: got healthy result from http://10.187.24.132:4001
Member 5780625be722ce57 is healthy: got healthy result from http://10.187.24.134:4001
Member ec2aa2bbe0c891b7 is healthy: got healthy result from http://10.187.27.200:4001
For the new node, first hear the etcd, delete the directory, then restart, and then join the cluster after reboot.
In addition, we deploy etcd through K8s, and other methods are not suitable. In addition, we highly recommend using K8s for deployment.