On the Use of Convolutional Neural Networks in the Tasks of Assessing the Security of Speech Acoustic Information
摘要
The paper focuses on the analysis of various architectures of convolutional neural networks and their application in the context of speech intelligibility. Four common architectures have been studied: VGGNet, GoogLeNet, AlexNet and ResNet. In the study, ResNet and AlexNet were tested, and a typical convolutional neural network model was created for comparison. For verification, a data set was compiled containing spectrograms of noisy audio recordings of speech, classified according to seven levels of intelligibility from 20% to 80%. All the studied architectures have been successfully trained, but special attention is paid to the ResNet architecture, since it has achieved high recognition accuracy with relatively low losses in the learning process. The study of convolutional neural network architectures and their practical application make it possible to expand the possibilities of assessing the security of speech acoustic information.