错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HTCondor Cluster Monitoring

  • E. Tsamtsurov,
  • N. Balashov

摘要

Abstract

As part of participation in various experiments, JINR provides computing resources in the form of a Batch cluster deployed as virtual machines in the JINR cloud based on the HTCondor system. Since a Batch processing system is a multi-component complex system, one of the key aspects of ensuring its smooth operation is constant monitoring of the state of its main components. The paper presents the developed HTCondor cluster monitoring system based on the Node Exporter, Prometheus, and Grafana technology stack. The general structure of the monitoring system is considered, the processes occurring in it are described. Developments are open and published, which allows them to be freely integrated into third-party infrastructures.