An Empirical Review of the Effectiveness of Different Language Processing Approaches in Software Code Vulnerability Detection
摘要
The advancement of software vulnerability detection tools has accelerated in recent years. However, the prevalence and severity of vulnerabilities continue to escalate, posing significant threats to computer security and information safety. Numerous detection methodologies have been proposed to address this, with machine learning-based approaches demonstrating notable promise. In this paper, we present a comprehensive review of state-of-the-art (SOTA) architectures that leverage Deep Learning (DL) and Natural Language Processing (NLP) or Large Language Models (LLMs) for identifying vulnerabilities. We systematically examine the efficiency of these cutting-edge architectures and performance analysis. We aim to uncover novel approaches for maximizing the potential of existing architectures to enhance vulnerability detection. During our research, we identified three key research questions: effective integration of NLP and DL technologies, strengths and limitations of LLMs in this domain, and comparative analysis of LLMs versus integrated NLP-DL approaches. In addition, we discuss the challenges and experimental constraints encountered in this domain, offering insights into future research directions. This study aims to inspire further exploration of innovative methodologies and contribute to the development of more robust cybersecurity solutions.