Addressing the limitations of Optical Mark Recognition (OMR) in grading Scantron 882-E answer sheets, this work explores the use of artificial intelligence (AI) to provide a more effective solution. It is a daunting and tedious task to curate large datasets for training deep learning (DL) models. To overcome this, synthetic datasets have become a popular alternative. This work examines the potential benefits of leveraging deep learning models, specifically generative adversarial networks (GANs), to generate large, diverse, and realistic images from a small, manually created dataset of pencil markings. Specifically, the approach focuses on generating synthetic data by modifying contours of collected markings, and flipping the markings on their x- and y- axes. The generated data is then used to train a CNN for determining answers on Scantron answer sheets. This approach demonstrated a significant reduction in the time required to acquire the extensive number of images needed for successfully training DL models and successfully determined the selected answers on Scantron sheets with 100% accuracy given nine real world samples.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generating Synthetic Datasets to Improve Deep Learning Model Training

  • Kolten Pulliam,
  • Terry Griffin,
  • Catherine Stringfellow

摘要

Addressing the limitations of Optical Mark Recognition (OMR) in grading Scantron 882-E answer sheets, this work explores the use of artificial intelligence (AI) to provide a more effective solution. It is a daunting and tedious task to curate large datasets for training deep learning (DL) models. To overcome this, synthetic datasets have become a popular alternative. This work examines the potential benefits of leveraging deep learning models, specifically generative adversarial networks (GANs), to generate large, diverse, and realistic images from a small, manually created dataset of pencil markings. Specifically, the approach focuses on generating synthetic data by modifying contours of collected markings, and flipping the markings on their x- and y- axes. The generated data is then used to train a CNN for determining answers on Scantron answer sheets. This approach demonstrated a significant reduction in the time required to acquire the extensive number of images needed for successfully training DL models and successfully determined the selected answers on Scantron sheets with 100% accuracy given nine real world samples.