错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Classification of Images Extracted from Scientific Documents for Cyber Deception

  • Ghanshyam S. Bopche,
  • Saloni Pawar,
  • Nilin Prabhaker

摘要

Protection of scientific documents from unauthorized access is crucial as they usually contain mission or business-critical information such as proprietary research data, innovative ideas, novel discoveries, data about industry collaboration and commercial interests, etc. Existing security controls are insufficient for protecting such sensitive documents from sophisticated cyberattacks such as Advanced Persistent Threats (APTs). Recent security solutions focus on data-level Cyber deception wherein multiple believable fake versions of intellectual property (IP) documents were generated and deployed throughout the enterprise network to slow down the adversary who needs to correctly identify the legitimate document hidden among the set of fake documents. As an integral component of scientific documents, images or figures convey critical information and complement textual content. Therefore, scientific images must also be faked while generating believable fake documents. These images may be different types but are not limited to diagrams, schematics, graphs, charts, simulation outputs, plots, flowcharts, and medical illustrations. These images need to be accurately classified before creating their believable fakes. However, the diversity and complexity of scientific images or charts complicate their accurate classification. This paper has tested several image classification models, such as SVM, Decision Tree, Random Forest, CNN, VGG16, InceptionV3, ResNet-50, and ResNet-101, to classify scientific images extracted from technical scientific documents. We have chosen DocFigure - a benchmark dataset of scientific annotated images for the training and testing of selected models. Our experiment illustrates that ResNet-101 is suitable for classifying scientific images or charts.