Umplc: the first longitudinal learner corpus of Portuguese
摘要
This paper introduces the University of Macau Portuguese Learner Corpus (UMPLC), an annotated longitudinal learner corpus of Portuguese. The corpus comprises 933 compositions (totaling 209,097 tokens) produced by 121 Chinese undergraduate students over six consecutive semesters (from their first to third year of study). At the moment, three text versions are publicly available: a transcribed version, a spell-checked version, and an annotated version with part of speech and lemma tags. We also provide an online search interface as a convenient tool for exploring the corpus. The UMPLC is a valuable language resource for researchers and educators. It enables a wide range of linguistic studies, and contributes to the development of pedagogical applications in the context of teaching and learning of Portuguese as a second language.