HTCondor Cluster Monitoring
摘要
Abstract
As part of participation in various experiments, JINR provides computing resources in the form of a Batch cluster deployed as virtual machines in the JINR cloud based on the HTCondor system. Since a Batch processing system is a multi-component complex system, one of the key aspects of ensuring its smooth operation is constant monitoring of the state of its main components. The paper presents the developed HTCondor cluster monitoring system based on the Node Exporter, Prometheus, and Grafana technology stack. The general structure of the monitoring system is considered, the processes occurring in it are described. Developments are open and published, which allows them to be freely integrated into third-party infrastructures.