This study presents a novel machine learning system that focuses on the analysis and detection of malware by leveraging intrinsic file features. The present study involves the analysis of a comprehensive dataset comprising features extracted from both known harmful and benign files. The ensemble Extra Trees Classifier algorithm is employed for the purpose of feature selection. The chosen features are utilized for the purpose of training and assessing Random Forest and Gradient Boosting classifiers, which are known for their ability to efficiently handle intricate feature interdependencies. Extensive performance evaluations conducted via cross-validation and testing provide empirical evidence that machine learning algorithms may effectively differentiate between malicious and benign files by leveraging intrinsic characteristics, resulting in an accuracy rate of over 95%. The importance of feature selection becomes evident in its role in improving generalization performance. Both classifiers exhibit strong and consistent performance. This study represents a notable advancement in the automation of malware analysis through the utilization of artificial intelligence. The proposed refinements have the ability to enhance existing signature-based methods and bolster defensive measures against the ever-changing landscape of malware threats.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Malware Analysis Using AI

  • Taher Kutbuddin Kapadia,
  • Vedangi Nilesh Gupte,
  • Bhaumik Hitesh Thakker,
  • Tushar Vitthal Sawant

摘要

This study presents a novel machine learning system that focuses on the analysis and detection of malware by leveraging intrinsic file features. The present study involves the analysis of a comprehensive dataset comprising features extracted from both known harmful and benign files. The ensemble Extra Trees Classifier algorithm is employed for the purpose of feature selection. The chosen features are utilized for the purpose of training and assessing Random Forest and Gradient Boosting classifiers, which are known for their ability to efficiently handle intricate feature interdependencies. Extensive performance evaluations conducted via cross-validation and testing provide empirical evidence that machine learning algorithms may effectively differentiate between malicious and benign files by leveraging intrinsic characteristics, resulting in an accuracy rate of over 95%. The importance of feature selection becomes evident in its role in improving generalization performance. Both classifiers exhibit strong and consistent performance. This study represents a notable advancement in the automation of malware analysis through the utilization of artificial intelligence. The proposed refinements have the ability to enhance existing signature-based methods and bolster defensive measures against the ever-changing landscape of malware threats.