In-Network Monitoring Strategies for HPC Cloud
摘要
The optimized network architectures and interconnect technologies employed in high-performance cloud computing environments introduce challenges when it comes to developing monitoring solutions that effectively capture relevant network metrics. Moreover, network monitoring often involves capturing and analyzing a large volume of network traffic data. This process can introduce additional overhead and consume system resources, potentially impacting the overall performance of HPC applications. Balancing the need for monitoring with minimal disruption to application performance is a key challenge. In this paper, we study different strategies to enable a low-overhead monitoring system utilizing emerging programmable network devices.