A Framework Using Feature Maps Improving Speech
摘要
Proper speech helps preschoolers’ cognitive and social development. A hybrid model including Wav2Vec 2.0 for audio feature extraction, Bi-LSTM, and an attention mechanism improves preschoolers’ speech competency categorization. LSTM, Bi-LSTM with Attention Mechanism, and Proposed Bi-LSTM were examined. All tests confirm the Proposed model’s 98.26% accuracy, precision, recall, and F1 score. Confusion matrix: Ordinary 98.55%, Incompetent 100%, Proficient 96.0%. The recommended model has a 0.98 ROC AUC, compared to 0.95 for Bi-LSTM and 0.90 for LSTM, suggesting excellent class difference. Adding attention processes to Bi-LSTM enhances toddler linguistic proficiency. For the Proposed approach, speech delivery classification efficiency increases early childhood speech therapy.