Voice Liveness Detection Using Bump Wavelet with CNN
摘要
In this work, Continuous Wavelet Transform (CWT) based features are proposed for Voice Liveness Detection (VLD). In particular, bump wavelet features are extracted from raw speech and classified using our proposed CNN architecture for VLD task. Determining whether a speech is coming from a live speaker enables us in developing robust countermeasures against spoofing attacks on Automatic Speaker Verification (ASV) systems, specially against replay attacks. Liveness of a speaker is attributed by the pop noise present in the live (genuine) speech signal. We utilize this characteristic of live speech for VLD task. We perform VLD using our approach and compare the experimental results with two existing approaches: 1) STFT-based baseline approach, 2) CQT-based approach, which gave accuracy as \(62.08\%\) and \(66.49\%\) , respectively. On the other hand, we achieve a highly improved accuracy of \(80.19\%\) . Furthermore, we also perform phoneme-based analysis and achieve higher accuracy values for all the phoneme types when compared with the two existing approaches.