Log real-time analysis architecture
When working in the last company, the design of log collection and real-time analysis framework is relatively simple:
Flume-ng + rocketmq + storm + redis + front-end display
In the message queue part, we started with kafka, but kafka is weak in supporting backtracking consumption and repeated consumption, and is also relatively weak in terms of data security. Later, we changed it to Ali's rocketmq.
Considering that the amount of data we have is not very large, it is enough to support it, but at the rocketmq layer, messages sometimes accumulate because of abnormal network problems, resulting in message queues being flooded, and the stability is not very high. Later, I consulted colleagues from other departments. Their approach is to add an additional layer of mongodb at the message queue level, and the message queue layer only retains the index information of the messages. The entity information of the message is saved in mongodb, which can avoid this problem. Later, due to various reasons, we did not try this method again.
Other common scenarios:
Logstash + elasticsearch + kibana
Fluentd + influxdb + grafana
Flume-ng + kafka + storm
Kafka + spark streaming + redis