Toward automating the rhythmic analysis of speech: a comparative study of English spoken by American and Thai speakers
摘要
Efforts to classify language rhythm have significantly contributed to methodologies for identifying acoustic features that define speech “beats”. This study advances the field through three key contributions: (1) validation of algorithmic rhythm detection, (2) characterization of distinctive rhythmic features in native versus non-native speech, and (3) development of an automated, scalable approach for rhythmic analysis. The effectiveness of the maxD parameter—representing the moment of fastest energy increase—was evaluated as an alternative to manual annotation for identifying syllabic beats. Speech samples from 34 speakers (17 American and 17 Thai) were analyzed using statistical and machine learning models, including Support Vector Machine, Random Forest, Gradient Boosting, and Logistic Regression, to classify rhythmic patterns. Results indicate that the maxD parameter demonstrates strong alignment with manually annotated beat locations, achieving high predictive accuracy for native speakers (RMSE: 0.1182, MAPE: 0.3290%). Additionally, SHapley Additive exPlanations (SHAP) analysis revealed key rhythmic features distinguishing the two groups: maxD_value, intensity_ratio, and max_intensity were identified as key characteristics of native speech, while shorter vowel and syllable durations were established as distinctive native features. Systematic differences in maxD, pitch, and intensity alignment patterns between native and non-native speech were documented, particularly in content words and multisyllabic structures. These findings contribute to the ultimate goal of improving non-native speech by identifying critical rhythmic features that differentiate native and non-native speakers. This research study lays the groundwork for targeted pronunciation training tools and language-learning applications, offering a foundation for enhancing naturalness and fluency in second language acquisition.