Quantifying job-level carbon efficiency in HPC: an empirical study based on the PM100 dataset
摘要
This study utilizes the PM100 dataset to quantitatively estimate job-level carbon emissions and analyze efficiency across different resource configurations. A multilayer perceptron (MLP) regression model was applied to predict emissions using execution time along with CPU, memory, and node-level power consumption data. To evaluate efficiency, we proposed the Carbon Efficiency Score (CES), which enables the classification of jobs into efficiency tiers. The analysis revealed that long-running jobs with excessive memory usage tend to exhibit low efficiency, whereas jobs with balanced resource configurations demonstrate relatively higher efficiency. CES-based classification further showed a difference of more than 200-fold between the most and least efficient jobs. Overall, this study provides a foundational framework for developing carbon-aware scheduling strategies in HPC environments and offers practical insights for the design of sustainable supercomputing operational policies.