As a key feature for teachers’ skills development in the Merdeka Mengajar Platform, Pelatihan Mandiri faced scaling challenges due to the exponential growth of users and submitted reporting documents (named: Aksi Nyata). These documents require validation for users to receive formal certificates from the Ministry of Education. Manual checking by human validators proved unsustainable with the rapid increase in document submissions. To overcome with the challenges, an automated validation system was initiated. This system employs a two-stage process as follows. The first stage utilizes distinct components of the Aksi Nyata submissions to filter out plagiarized documents using the Levenshtein distance method and addresses common errors, such as missing reflections and image documentation, through a rule-based approach. In the second stage, the entire content of Aksi Nyata submissions is preprocessed through tokenization, stopword removal, and word embeddings. State-of-the-art machine learning models were then trained on manually validated data sets, selecting those with the highest coverage and lowest error rates. To date, the model has processed up to 80% of document submissions per month with a false positive rate of 2%. This system has also resulted in speeding up the validation process time by 76%, henceforth accelerating the feedback process for users.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Automatic Document Validation System in the Platform Merdeka Mengajar by Using Word Embeddings and Supervised Classification Models

  • Figarri Keisha,
  • Andika R. Hakim,
  • Bagoes R. Widiarso,
  • Raden Z. Hermawan,
  • Aghnia M. Safira,
  • Rizka Azmira,
  • Putri W. Novianti

摘要

As a key feature for teachers’ skills development in the Merdeka Mengajar Platform, Pelatihan Mandiri faced scaling challenges due to the exponential growth of users and submitted reporting documents (named: Aksi Nyata). These documents require validation for users to receive formal certificates from the Ministry of Education. Manual checking by human validators proved unsustainable with the rapid increase in document submissions. To overcome with the challenges, an automated validation system was initiated. This system employs a two-stage process as follows. The first stage utilizes distinct components of the Aksi Nyata submissions to filter out plagiarized documents using the Levenshtein distance method and addresses common errors, such as missing reflections and image documentation, through a rule-based approach. In the second stage, the entire content of Aksi Nyata submissions is preprocessed through tokenization, stopword removal, and word embeddings. State-of-the-art machine learning models were then trained on manually validated data sets, selecting those with the highest coverage and lowest error rates. To date, the model has processed up to 80% of document submissions per month with a false positive rate of 2%. This system has also resulted in speeding up the validation process time by 76%, henceforth accelerating the feedback process for users.