Analysis of audio evidence is a major component of forensic casework. Cases usually deal with threat calls, phishing scams, etc. which lead to considerable emotional distress and financial losses. Forensic sciences have traditionally dealt with audio evidence in three directions of enquiry: determining authenticity, enhancing audio recordings for intelligibility and audibility, and interpretation of sonic evidence. With the advent of Generative AI (GenAI), the occurrence of criminal activities using Deepfake audio has become easier. It can take the form of voice phishing; spreading of misinformation; impersonation for blackmail, harassment, and fraudulent business calls; violation of privacy; and creation of fake news to cause unrest which also leads to an erosion of trust in media, fake emergency calls, and manipulation of audio evidence in a legal context. We review contemporary research and established methods of audio forensics in this paper. The generally accepted techniques of audio forensics involve critical listening of the evidence and known samples, based on linguistics, phonetics, and speaker characteristics. The waveform is examined to look for signs of tampering and aural-spectrographic examination is carried out for comparing the spectrogram of the evidence with those of the samples recorded of the suspect based on patterns of spectral features, formant shapes, etc. of the voices. To study the state of the art in AI-generated speech technologies, recent work on synthetic speech generation is reviewed. Audio generation algorithms can be text-to-speech (TTS) and voice conversion (VC). Rapid development of Generative AI through neural networks, Natural Language Processing (NLP), and attention systems using Transformers has allowed the generation of high-quality realistic voices. Recent work on detection of Deepfakes has also been studied to identify parameters such as cepstral coefficients and bispectral analysis that can be utilized to develop a forensic method for identifying Deepfakes. We present a critical analysis of the implications of the occurrence of Deepfake audio in audio forensics to identify gaps in research and explore new directions in which methods of analysis of audio evidence in forensic laboratories should develop to address the emerging concerns with AI-generated media.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Future of Audio Forensics: Exploring the Effect of Generative AI

  • Soham Gangopadhyay,
  • Priyanka Gopi,
  • Prateek Pandya,
  • Ashish Mani,
  • Sumit Goswami

摘要

Analysis of audio evidence is a major component of forensic casework. Cases usually deal with threat calls, phishing scams, etc. which lead to considerable emotional distress and financial losses. Forensic sciences have traditionally dealt with audio evidence in three directions of enquiry: determining authenticity, enhancing audio recordings for intelligibility and audibility, and interpretation of sonic evidence. With the advent of Generative AI (GenAI), the occurrence of criminal activities using Deepfake audio has become easier. It can take the form of voice phishing; spreading of misinformation; impersonation for blackmail, harassment, and fraudulent business calls; violation of privacy; and creation of fake news to cause unrest which also leads to an erosion of trust in media, fake emergency calls, and manipulation of audio evidence in a legal context. We review contemporary research and established methods of audio forensics in this paper. The generally accepted techniques of audio forensics involve critical listening of the evidence and known samples, based on linguistics, phonetics, and speaker characteristics. The waveform is examined to look for signs of tampering and aural-spectrographic examination is carried out for comparing the spectrogram of the evidence with those of the samples recorded of the suspect based on patterns of spectral features, formant shapes, etc. of the voices. To study the state of the art in AI-generated speech technologies, recent work on synthetic speech generation is reviewed. Audio generation algorithms can be text-to-speech (TTS) and voice conversion (VC). Rapid development of Generative AI through neural networks, Natural Language Processing (NLP), and attention systems using Transformers has allowed the generation of high-quality realistic voices. Recent work on detection of Deepfakes has also been studied to identify parameters such as cepstral coefficients and bispectral analysis that can be utilized to develop a forensic method for identifying Deepfakes. We present a critical analysis of the implications of the occurrence of Deepfake audio in audio forensics to identify gaps in research and explore new directions in which methods of analysis of audio evidence in forensic laboratories should develop to address the emerging concerns with AI-generated media.