Hybrid Approach to HPC Cluster Telemetry and Hardware Log Analytics

Sep 1, 2020·

Justin Thaler

Woong Shin

Steven Roberts

James H. Rogers

Todd Rosedahl

· 0 min read

Cite DOI URL

Abstract

The number of computer processing nodes and processor cores in cluster systems is growing rapidly. Discovering, and reacting to, a hardware or environmental issue in a timely manner enables proper fault isolation, improves quality of service, and improves system up-time. In the case of performance impacts and node outages, RAS policies can direct actions such as job quiescence or migration. Additionally, power consumption, thermal information, and utilization metrics can be used to provide cluster energy and cooling efficiency improvements as well as optimized job placement. This paper describes a highly scalable telemetry architecture that allows event aggregation, application of RAS policies, and provides the ability for cluster control system feedback. The architecture advances existing approaches by including both programmable policies, which are applied as events stream through the hierarchical network to persistence storage, and treatment of sensor telemetry in an extensible framework. This implementation has proven robust and is in use in both cloud and HPC environments including the Summit system of 4,608 nodes at Oak Ridge National Laboratory.

Type

Conference

Publication

2020 IEEE High Performance Extreme Computing Conference (HPEC)

Last updated on Sep 1, 2020

Monitoring Telemetry High Performance Computing Summit Supercomputer

Authors

Woong Shin

HPC Systems/Software/Data Engineer, Computer Systems Researcher

← Global Experiences with HPC Operational Data Measurement, Collection and Analysis Sep 1, 2020

Data Jockey: Automatic Data Management for HPC Multi-tiered Storage Systems May 1, 2019 →