Machine Learning Based Malware Identification and Classification in PDF: A Review Paper
摘要
Today’s modern antivirus software fails to provide protection against malicious PDF (Portable Document Format) files, which is considered a threat to system security. This study introduces a new machine-learning-based classification approach to PDF malware to reduce PDF malware somewhat. A unique feature of this system is static and dynamic checking of provided PDF files. As a result, each of these classification algorithms may more accurately determine the right document type, negative rate (FNR), and F1 score, which is the most appropriate strategy evaluated. The proposed system is also subject to malicious attacks that obfuscate the malicious code in the PDF file by hiding it from the PDF parser during the processing stage. It was determined that the proposed technique outperformed the prior art, which had an F1 metric of 0.978, by achieving an F1 metric of 0.986 utilizing a random forest (RF) classifier. Compared to current solutions, this is very effective at detecting malware encoded in PDF files.