<p>Among Igbo native speakers of English, a speech disorder called lambdacism occurs, in which/l/l/ is articulated as /r/. This study presents SVM-Lambda-Tune, a proposed machine-learning-based model for automated identification and correction of pronunciation errors. It classifies the /l/ phoneme as correct or incorrect using an optimized Support Vector Machine (SVM) by tuning it with Principal Component Analysis (PCA) and Recursive Feature Elimination (RFE). The model was trained on a dataset of audio files containing 96 recorded English words, each with correct and incorrect pronunciations, and appropriately pre-processed, from which features such as Mel-Frequency Cepstral Coefficients (MFCC) and Term Frequency-Inverse Document Frequency (TF-IDF) were extracted. At the same time, phoneme alignment was evaluated using Dynamic Time Warping (DTW) to boost its usability. A Multimodal Corrective Feedback Module (MCFM) was applied for the integration of real-time feedback and the expansion of the model’s applicability to a broader range of phoneme error patterns and to reliably detect mispronunciations and suggest steps for proper articulation through guided audio-visual feedback. The model also achieved 95% classification accuracy, demonstrating the effectiveness of feature selection techniques for classification and dimensionality reduction, with a 16.32% reduction in Word Error Rate (WER). This provides a scalable and accessible solution for pronunciation problems and serves as a foundation for further developments in speech technology, learning aids for people with disabilities, and forensic phonetics.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automated lambdacism correction using machine learning for rural speech therapy

  • Vivian Chinwe Okanme,
  • Rajesh Prasad,
  • Francisca Nonyelum Ogwueleka,
  • Fatimah Binta Abdullahi

摘要

Among Igbo native speakers of English, a speech disorder called lambdacism occurs, in which/l/l/ is articulated as /r/. This study presents SVM-Lambda-Tune, a proposed machine-learning-based model for automated identification and correction of pronunciation errors. It classifies the /l/ phoneme as correct or incorrect using an optimized Support Vector Machine (SVM) by tuning it with Principal Component Analysis (PCA) and Recursive Feature Elimination (RFE). The model was trained on a dataset of audio files containing 96 recorded English words, each with correct and incorrect pronunciations, and appropriately pre-processed, from which features such as Mel-Frequency Cepstral Coefficients (MFCC) and Term Frequency-Inverse Document Frequency (TF-IDF) were extracted. At the same time, phoneme alignment was evaluated using Dynamic Time Warping (DTW) to boost its usability. A Multimodal Corrective Feedback Module (MCFM) was applied for the integration of real-time feedback and the expansion of the model’s applicability to a broader range of phoneme error patterns and to reliably detect mispronunciations and suggest steps for proper articulation through guided audio-visual feedback. The model also achieved 95% classification accuracy, demonstrating the effectiveness of feature selection techniques for classification and dimensionality reduction, with a 16.32% reduction in Word Error Rate (WER). This provides a scalable and accessible solution for pronunciation problems and serves as a foundation for further developments in speech technology, learning aids for people with disabilities, and forensic phonetics.