Skip to main content
SHARE
Publication

Pattern-based Modeling of Multiresilience Solutions for High-Performance Computing...

by Rizwan A Ashraf, Saurabh Hukerikar, Christian Engelmann
Publication Type
Conference Paper
Book Title
Proceedings of the 2018 ACM/SPEC International Conference on Performance Engineering
Publication Date
Page Numbers
80 to 87
Publisher Location
New York, New York, United States of America
Conference Name
9th ACM/SPEC International Conference on Performance Engineering (ICPE 2018)
Conference Location
Berlin, Germany
Conference Sponsor
Association for Computing Machinery (ACM) and Standard Performance Evaluation Corporation (SPEC)
Conference Date
-

Resiliency is the ability of large-scale high-performance computing (HPC) applications to gracefully handle errors, and recover from failures. In this paper, we propose a pattern-based approach to constructing resilience solutions that handle multiple error modes. Using resilience patterns, we evaluate the performance and reliability characteristics of detection, containment and mitigation techniques for transient errors that cause silent data corruptions and techniques for fail-stop errors that result in process failures. We demonstrate the design and implementation of the multiresilience solution based on patterns instantiated across multiple layers of the system stack. The patterns are integrated to work together to achieve resiliency to different error types in a performance-efficient manner.