Parallelized Data Replication of Multi-Petabyte Storage Systems
SessionHPCSYSPROS20
Event Type
Workshop
W
TimeFriday, 13 November 202012:10pm - 12:30pm EDT
LocationTrack 4
DescriptionThis paper presents the architecture of a highly parallelized data replication workflow implemented at The University of Sydney that forms the disaster recovery strategy for two 8-petabyte research data storage systems at the University. The solution leverages DDN’s GRIDScaler appliances, the information lifecycle management feature of the IBM Spectrum Scale File System, rsync, GNU Parallel and the MPI dsync tool from mpiFileUtils. It achieves high performance asynchronous data replication between two storage systems at sites 40km apart. In this paper, the methodology, performance benchmarks, technical challenges encountered and fine-tuning improvements in the implementation are presented.