Through child’s eyes: evaluating LLMs for complex word identification in children’s stories
摘要
The complexity of text significantly impacts reading comprehension, especially for young readers. Identifying complex vocabulary and then replacing it with a simple alternative is essential for improving reading accessibility among children. This study investigates the effectiveness of large language models (LLMs) to perform complex word identification (CWI) within children’s stories through a zero-shot prompting approach. We designed a child-centric annotation interface to create ground truth by learners aged 8–10, capturing judgments at the document level rather than isolated sentences. To ensure explainability, we conducted statistical analyses of linguistic features across complex and simple words, revealing how structural features impact model decisions. We evaluate four LLMs, GPT-4o and open-source models Llama3, Mistral, and Gemma2, against these annotations. Results show that GPT-4o demonstrates better alignment with children’s perception of word complexity, achieving a balanced trade-off between recall and precision. These findings establish the effectiveness of LLMs and emphasize the value of considering target user perception to evaluate the model’s decisions.