Corpus-Based Analysis of Lexical Features of Mongolian Language Policy Text
摘要
Like other policy texts, language policy texts also need policy text analysis. Leveraging a corpus of 100 policy documents, this study investigates various linguistic attributes, including the distribution of parts of speech, type-token ratio, and lexical density, through data comparative analysis. Furthermore, the paper categorizes the corpus into twelve distinct types, encompassing instructions, decisions, notices, reports, regulations, ways, rules, methods, summaries, plans, speeches and papers. Employing natural language processing techniques, the study also utilizes frequency statistics and wordclouds to provide both word frequency statistical tables and visual wordcloud representations of the Mongolian language policy text corpus.