<p>Reference-free audio quality assessment is a valuable tool in many areas, such as audio recordings, vinyl production, and communication systems. Therefore, evaluating the reliability and performance of such tools is crucial. This paper builds on previous research by analyzing the performance of four additional algorithms in detecting perceptible impulsive noise (clicks) based on auditory models. We compared the results of eight algorithms, hypothesizing that computationally simpler algorithms could perform as well as more complex ones. We obtained a set of audio signals, with and without clicks, annotated by human subjects from a publicly available dataset. Audio signal sets are categorized based on the obtained annotation results to train the algorithms for different levels of the experiments. Experiments containing cross-validation are done for multiple parameters of algorithms. The algorithm training is based on maximizing a discriminability metric (<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13636_2024_389_Article_IEq1.gif" Format="GIF" Height="15" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(A'\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>A</mi> <mo>′</mo> </msup> </math></EquationSource> </InlineEquation>). Evaluation criteria of the algorithms included the hit rate, false alarm rate, <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13636_2024_389_Article_IEq2.gif" Format="GIF" Height="15" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(A'\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>A</mi> <mo>′</mo> </msup> </math></EquationSource> </InlineEquation>, and computational time. Our findings indicate that computationally simpler auditory models have performed as well as computationally more complex ones, while conventional models exhibit lower performance. Conclusively, the ERBlet transform based algorithm demonstrated superior performance in terms of <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13636_2024_389_Article_IEq3.gif" Format="GIF" Height="15" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(A'\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>A</mi> <mo>′</mo> </msup> </math></EquationSource> </InlineEquation> and robustness. This paper provides insights into the capabilities of auditory models in a practical use case of perceptible click detection. The results presented here can help research and develop such algorithms for vinyl production, audio archiving, podcasting, music production, and telecommunications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Performance evaluation of perceptible impulsive noise detection methods based on auditory models

  • Arda Özdoğru,
  • František Rund,
  • Karel Fliegel

摘要

Reference-free audio quality assessment is a valuable tool in many areas, such as audio recordings, vinyl production, and communication systems. Therefore, evaluating the reliability and performance of such tools is crucial. This paper builds on previous research by analyzing the performance of four additional algorithms in detecting perceptible impulsive noise (clicks) based on auditory models. We compared the results of eight algorithms, hypothesizing that computationally simpler algorithms could perform as well as more complex ones. We obtained a set of audio signals, with and without clicks, annotated by human subjects from a publicly available dataset. Audio signal sets are categorized based on the obtained annotation results to train the algorithms for different levels of the experiments. Experiments containing cross-validation are done for multiple parameters of algorithms. The algorithm training is based on maximizing a discriminability metric ( \(A'\) A ). Evaluation criteria of the algorithms included the hit rate, false alarm rate, \(A'\) A , and computational time. Our findings indicate that computationally simpler auditory models have performed as well as computationally more complex ones, while conventional models exhibit lower performance. Conclusively, the ERBlet transform based algorithm demonstrated superior performance in terms of \(A'\) A and robustness. This paper provides insights into the capabilities of auditory models in a practical use case of perceptible click detection. The results presented here can help research and develop such algorithms for vinyl production, audio archiving, podcasting, music production, and telecommunications.