The exponential growth in the usage of network technologies has led to security attacks and threats. Software vulnerability is one of the major reasons for the breaching of such attacks by attackers. Various methods including static analysis, dynamic analysis, as well as hybrid analysis are intended to predict the vulnerabilities in software, which have some drawbacks related to it. Trending machine learning algorithms have played a crucial role in identifying the key aspects of the weaknesses or flaws leading to vulnerable statistics, with the help of finding unusual patterns or behavior inside the system or file. This paper contributes to identifying the vulnerabilities in the C source code with the help of the Term Frequency-Inverse Document Frequency (TFID) as the embedding technique for vectorization as well as feature extraction, and machine learning algorithms including, Random Forest (RF), Support Vector Machine (SVM), K-nearest neighbor (KNN), Logistic Regression (LR), for prediction. The performance analysis of the machine learning techniques is done on a dataset named LibPNG.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Study of Machine Learning Algorithms for Predicting Vulnerability in Software

  • Nitika,
  • Kuldeep Kumar

摘要

The exponential growth in the usage of network technologies has led to security attacks and threats. Software vulnerability is one of the major reasons for the breaching of such attacks by attackers. Various methods including static analysis, dynamic analysis, as well as hybrid analysis are intended to predict the vulnerabilities in software, which have some drawbacks related to it. Trending machine learning algorithms have played a crucial role in identifying the key aspects of the weaknesses or flaws leading to vulnerable statistics, with the help of finding unusual patterns or behavior inside the system or file. This paper contributes to identifying the vulnerabilities in the C source code with the help of the Term Frequency-Inverse Document Frequency (TFID) as the embedding technique for vectorization as well as feature extraction, and machine learning algorithms including, Random Forest (RF), Support Vector Machine (SVM), K-nearest neighbor (KNN), Logistic Regression (LR), for prediction. The performance analysis of the machine learning techniques is done on a dataset named LibPNG.