Developing Multilingual Glossaries for STEM Terminology Using AI-NLP
摘要
Educating Indigenous populations in science, technology, engineering, and mathematics (STEM) remains an urgent necessity in India. Many students from Indigenous populations are first-generation learners, which results in gaps in domain knowledge at all educational levels. STEM education is further complicated by the fact that most students from Indigenous communities are not conversant with English, the language in which most scientific books, journals, and articles are written in. Students are unable to relate with many subject-specific terms, which results in impaired learning outcomes. Developing mother-tongue-based multilingual glossaries in the students’ own languages may be one solution. However, such glossaries are best built at the grassroots level, with input from practitioners of Indigenous languages and a common consensus. The practitioners must also be able to correctly render the meaning of the technical term sufficiently in their Indigenous languages. The process of creating technical terms as teaching aids in Indigenous languages must thus be accompanied by explanations in a commonly understood language, and building an appropriate term using appropriate cognate root words. These curated root words can be used to create a corpus that can be fed to an NLP algorithm, which would synthesize word-equivalents in Indigenous languages. These new words for STEM terminology can then be validated by native speakers of Indigenous languages and tested in classrooms, with the acceptability, use, and validation forming the basis for a heuristic approach to improving AI-driven synthesis of STEM terminology word-equivalents.