<p>This paper introduces a noise-resilient method for segmenting Dravidian-accented Malayalam speech into syllable-like units. These sub-units, derived from acoustic cues that approximate syllables from a linguistic perspective, are referred to as syllable-like units. The proposed approach employs sonority estimation based on the Ramped Autocorrelation Coefficient (RAC), which effectively captures the auditory prominence of syllables even in adverse acoustic conditions. Unlike conventional segmentation techniques, the method exploits the noise-robust property of autocorrelation to detect sonority peaks and valleys, thereby improving segmentation reliability under background noise. The method is evaluated using the Dravidian Accented Malayalam Speech Database (DAMSD), which includes clean and simulated noisy speech generated with White Gaussian, Pink, Red, and Babble noise at three signal-to-noise ratio (SNR) levels: 20 dB, 10 dB, and 0 dB. The manually annotated syllable boundaries from the DAMSD corpus serve as the ground truth for evaluation. Experimental results demonstrate consistent segmentation performance across all noise conditions, with F1-scores ranging from 74.15% (clean) to 69.56% (Babble, 0 dB). The performance achieved is comparable to that reported for established text-independent and sonority-based syllable segmentation approaches, while additionally providing a systematic evaluation under multiple additive noise environments and SNR levels. The RAC-based method exhibits minimal degradation in performance at lower SNRs, indicating its suitability for robust syllable-like segmentation in adverse acoustic conditions. The proposed syllable-like units thus serve as stable, intermediate representations between phonemes and words, supporting noise-tolerant speech analysis and accent classification tasks. Overall, the study highlights the potential of noise-resilient syllable-like segmentation as a foundation for developing robust speech processing systems for under-resourced languages.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Noise-resilient syllable-like segmentation of Dravidian-Accented Malayalam speech using sonority estimation from ramped autocorrelation

  • Sunil John Eanthalamkuzhiyil,
  • Bibish Kumar K. T. Kokkoli Theruvath,
  • Muraleedharan K. M. Kulappirath Meethal,
  • Suni Kumar R. K. Rayaroth Kuniyil

摘要

This paper introduces a noise-resilient method for segmenting Dravidian-accented Malayalam speech into syllable-like units. These sub-units, derived from acoustic cues that approximate syllables from a linguistic perspective, are referred to as syllable-like units. The proposed approach employs sonority estimation based on the Ramped Autocorrelation Coefficient (RAC), which effectively captures the auditory prominence of syllables even in adverse acoustic conditions. Unlike conventional segmentation techniques, the method exploits the noise-robust property of autocorrelation to detect sonority peaks and valleys, thereby improving segmentation reliability under background noise. The method is evaluated using the Dravidian Accented Malayalam Speech Database (DAMSD), which includes clean and simulated noisy speech generated with White Gaussian, Pink, Red, and Babble noise at three signal-to-noise ratio (SNR) levels: 20 dB, 10 dB, and 0 dB. The manually annotated syllable boundaries from the DAMSD corpus serve as the ground truth for evaluation. Experimental results demonstrate consistent segmentation performance across all noise conditions, with F1-scores ranging from 74.15% (clean) to 69.56% (Babble, 0 dB). The performance achieved is comparable to that reported for established text-independent and sonority-based syllable segmentation approaches, while additionally providing a systematic evaluation under multiple additive noise environments and SNR levels. The RAC-based method exhibits minimal degradation in performance at lower SNRs, indicating its suitability for robust syllable-like segmentation in adverse acoustic conditions. The proposed syllable-like units thus serve as stable, intermediate representations between phonemes and words, supporting noise-tolerant speech analysis and accent classification tasks. Overall, the study highlights the potential of noise-resilient syllable-like segmentation as a foundation for developing robust speech processing systems for under-resourced languages.