Comparative Analysis of VGGish and YAMNet Models for Welding Defect Detection
摘要
Manufacturing, especially in welding, requires advanced tools to detect and classify sounds associated with the process that may lead to defects. This study investigates the applicability of pre-trained convolutional neural network models known as VGGish (Visual Geometry Group) and YAMNet (Yet Another Mixture-of-Experts Network) to analyze and improve welding defect detection. The methodology involved data collection and processing, as well as training and evaluation of the models in terms of accuracy and inference time. To determine the evaluation parameters, songs from two rock bands with similar styles were used to extract audio information, considering the audio amplitude, rhythm (pattern), and tempo (speed) of the songs. Additionally, the behavior of the audio was visualized using a spectrogram, which illustrates the variation of sound over time. The objective is to understand the acoustic structure of the songs, which can be useful for pattern detection and analysis of audio signals. It is expected that, after applying the pre-trained models VGGish and YAMNet in predicting audios from a custom model trained with these songs, one of the two pre-trained models will demonstrate superior performance in terms of accuracy and inference time. This denotes a pre-trained model with better probabilities, which can be further used in other research for welding audio, aiding in the development of the acoustic monitoring tool in manufacturing, especially within the welding process. This contributes to the efficient identification of defects and process optimization.