Kafka broker Leader-1 caused spark Streaming not to be consumed has been resolved
I. Description of the problem:
One machine cdh-003 in the Kafka production cluster hung up due to a physical failure, and the system failed to get up, making online spark Streaming real-time tasks unable to consume normally, and restarting real-time tasks failed. Check the status of kafka topic and find that broker Leader appears-1, as shown in the following figure
II. Problem Analysis
Kafka Broker Leader is-1, which means that a partition failed to elect Leader, so the real-time tasks consuming this Topic all have exceptions. After elimination, it is found that the suspended cdh-003 machine happens to be broker id 257. (But why wasn't 192 elected leader?)
Solution: Modify kafka metadata and manually specify kakfa Leader.
The kafka partition status information is stored on Zookeeper. My environment directory is/kafka/brokers/topics/. The specific operation is as follows:
1. Check the partition status of leader-1.
[zk: localhost:2181(CONNECTED) 2] get /kafka/brokers/topics/mds001/partitions/1/state
{"controller_epoch":87,"leader":-1,"version":1,"leader_epoch":96,"isr":[257]}
2. Forced to modify partition leader to 192
[zk: localhost:2181(CONNECTED) 3] set /kafka/brokers/topics/mds001/partitions/1/state {"controller_epoch":87,"leader":192,"version":1,"leader_epoch":96,"isr":[192]}
3. Check whether the modification is successful
[zk: localhost:2181(CONNECTED) 4] get /kafka/brokers/topics/mds001/partitions/1/state
{"controller_epoch":87,"leader":192,"version":1,"leader_epoch":96,"isr":[192]}
[zk: localhost:2181(CONNECTED) 5]
4. Restart Kafka service (must restart, I did not restart at the beginning, so SS consumption is still abnormal) 5. Restart Spark Streaming real-time task, at this time consumption is normal, it is perfectly solved