Models for Resilience Design Patterns
Event Type
Workshop
Extreme Scale Computing
Fault Tolerance
Reliability and Resiliency
W
TimeWednesday, 11 November 202011:05am - 11:35am EDT
LocationTrack 11
DescriptionResilience plays an important role in supercomputers by providing correct and efficient operation in case of faults, errors, and failures. Resilience design patterns offer blueprints for effectively applying resilience technologies. Prior work focused on developing initial efficiency and performance models for resilience design patterns. This paper extends it by (1) describing performance, reliability, and availability models for all structural resilience design patterns, (2) providing more detailed models that include flowcharts and state diagrams, and (3) introducing the Resilience Design Pattern Modeling (RDPM) tool that calculates and plots the performance, reliability, and availability metrics of individual patterns and pattern combinations.