Get the App
SLTechnology News&Howtos  ›  Servers  › 

The process of initial inspection and data recovery after server paralysis

Shulou Source: shulou.com Published: 2022-06-02 06:52:20 10月03日 Update

Server data recovery background

A server in a state-owned enterprise in Beijing suddenly crashed during normal operation. The server has 240 hard drives, including 24 hard drives for metadata storage, nine raid1 disk arrays and one raid10 disk array. The rest of the hard drives consist of an average of 36 raid5 disk arrays. The ultimate reason for this storage paralysis is that two hard disks in one of the disk arrays are offline one after another, thus affecting the unavailability of the entire server, so we have to contact Beijing data recovery Company for door-to-door detection and recovery of server data.

Initial check of server data recovery

The data recovery center arranges engineers to come to the customer site to conduct a simple initial inspection and evaluation of the failed server and then start data backup. Because offline hard disks belong to the same group of raid arrays, two different backup methods are adopted for servers, that is, full sector-level mirroring of offline raid and storage-level backup of other raid arrays without offline hard disks. In the backup process of the faulty raid array, it is found that there are a large number of irregular bad paths in one of the two disconnected hard drives, so the backup can not be carried out, so we have to replace and repair the firmware of the hard disk, but a large number of bad paths still exist.

Data analysis

The server data recovery engineer first makes a detailed analysis of the underlying structure of the faulty RAID array, and then virtual reassembles the raid array according to the analyzed raid information for further analysis. Through further analysis, it is found that the hard disk with a lot of bad channels is offline late, which may have a certain impact on the final data recovery results.

Log in to the management system of the storage device to obtain the basic information about the volume in the file system and find that there are two volumes in the file system, and then continue to analyze the directory and node information of the Meta volume and the indexing algorithm of the Meta volume to the Data volume.

Server data recovery

After the basic information necessary for data recovery is obtained through the analysis of the server data recovery engineer, the engineer writes a program to scan and parse the nodes and directory items, and derives the complete directory structure of the file system. Parse the pointer information in each node and record it in the database.

After random sampling testing of all the data recovered by the engineer, the customer confirms the integrity of the data and agrees to hand over the data recovery result, and the server data recovery is successful.

Tags: Data server service data recovery hard disk array analysis information backup engineering engineer disk disk array system storage failure file directory node process Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno macOS Shulou Technology OPPO Reno NVidia Shulou Information