From Failure to Insight: Analyzing Disk Breakdowns in Large-Scale HPC Environments

Disk failure data provides valuable insights for preventing failures, enhancing storage robustness, guiding system design and deployment, and ensuring reliable operations at data centers. This paper introduces two disk failure datasets collected from large-scale HPC production environments over the...

Celý popis

Uloženo v:

Podrobná bibliografie
Vydáno v:	SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis s. 484 - 495
Hlavní autoři:	George, Anjus, Wang, Meng, Hanley, Jesse, Ransom, Garrett Wilson, Bent, John, Zimmer, Christopher
Médium:	Konferenční příspěvek
Jazyk:	angličtina
Vydáno:	IEEE 17.11.2024
Témata:	Cause effect analysis Conferences Data centers Electric breakdown Failure data analysis Hard disk drives High performance computing HPC storage Production Reliability Reliability engineering Robustness Summit Supercomputer System analysis and design
On-line přístup:	Získat plný text
Tagy:	Přidat tag Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!

Popis
Shrnutí:	Disk failure data provides valuable insights for preventing failures, enhancing storage robustness, guiding system design and deployment, and ensuring reliable operations at data centers. This paper introduces two disk failure datasets collected from large-scale HPC production environments over the past five years, comprising over 5,000 failure records from more than 40,000 disks. We analyzed these datasets across multiple dimensions, including temporal, spatial, and relational trends, and performed a comprehensive reliability assessment. Our analysis yielded numerous observations and insights that influence various operational aspects of HPC storage systems. We believe this study offers a holistic understanding of disk failure trends likely to interest the HPC storage community.
DOI:	10.1109/SCW63240.2024.00070