<p>Static Android malware detection remains widely used because of its scalability and low analysis cost; however, reported performance can be inflated by duplicate-sensitive evaluation, repeated application artefacts, and insufficiently controlled data partitioning. This study presents a leakage-aware reassessment of two engineered static features, Behaviour Score (BS) and Obfuscation Level (OL), for Android malware detection. Starting from a 4547-sample dataset, duplicate cleaning reduced the corpus to 3863 applications by removing full duplicates and duplicate feature-label rows. Malware samples were collected from MalwareBazaar and the IEEE 2024 Maloid-DS labeled Android malware forensics dataset, while goodware samples were collected from the official F-Droid open-source Android application repository using an automated downloader. A group-aware 80/20 holdout split based on application identity was then applied, followed by five-fold cross-validation on the training partition. Four feature configurations were evaluated through controlled ablation using Logistic Regression, Support Vector Machine, Random Forest, Extra Trees, Multilayer Perceptron, and XGBoost. The best main holdout performance was obtained by XGBoost with the baseline+OL configuration, achieving 98.83% accuracy, 98.91% F1-score, 0.9981 ROC-AUC, 0.9987 PR-AUC, and 0.9766 MCC. In contrast, BS produced a smaller gain in raw predictive performance, while the combined BS+OL setting did not surpass OL alone on the main holdout test. Additional statistical robustness checks showed that the strongest configurations were closely clustered, indicating that the ablation differences should be interpreted cautiously rather than as large separations. Obfuscation-stratified analysis further showed consistently high performance across low-, medium-, and high-OL subsets, with only minor variation in F1-score. These findings indicate that, under leakage-aware evaluation, OL provides the strongest main holdout result in the current dataset, whereas BS remains useful as a compact and interpretable engineered abstraction. The study highlights the importance of leakage-aware evaluation in malware detection research and shows that engineered features should be assessed not only for accuracy, but also for empirical stability, interpretability, and robustness under stricter validation conditions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leakage-aware reassessment of behaviour score and obfuscation level for static Android malware detection

  • Ali Fenjan,
  • Mohammed Almulla,
  • Jalil Md. Desa

摘要

Static Android malware detection remains widely used because of its scalability and low analysis cost; however, reported performance can be inflated by duplicate-sensitive evaluation, repeated application artefacts, and insufficiently controlled data partitioning. This study presents a leakage-aware reassessment of two engineered static features, Behaviour Score (BS) and Obfuscation Level (OL), for Android malware detection. Starting from a 4547-sample dataset, duplicate cleaning reduced the corpus to 3863 applications by removing full duplicates and duplicate feature-label rows. Malware samples were collected from MalwareBazaar and the IEEE 2024 Maloid-DS labeled Android malware forensics dataset, while goodware samples were collected from the official F-Droid open-source Android application repository using an automated downloader. A group-aware 80/20 holdout split based on application identity was then applied, followed by five-fold cross-validation on the training partition. Four feature configurations were evaluated through controlled ablation using Logistic Regression, Support Vector Machine, Random Forest, Extra Trees, Multilayer Perceptron, and XGBoost. The best main holdout performance was obtained by XGBoost with the baseline+OL configuration, achieving 98.83% accuracy, 98.91% F1-score, 0.9981 ROC-AUC, 0.9987 PR-AUC, and 0.9766 MCC. In contrast, BS produced a smaller gain in raw predictive performance, while the combined BS+OL setting did not surpass OL alone on the main holdout test. Additional statistical robustness checks showed that the strongest configurations were closely clustered, indicating that the ablation differences should be interpreted cautiously rather than as large separations. Obfuscation-stratified analysis further showed consistently high performance across low-, medium-, and high-OL subsets, with only minor variation in F1-score. These findings indicate that, under leakage-aware evaluation, OL provides the strongest main holdout result in the current dataset, whereas BS remains useful as a compact and interpretable engineered abstraction. The study highlights the importance of leakage-aware evaluation in malware detection research and shows that engineered features should be assessed not only for accuracy, but also for empirical stability, interpretability, and robustness under stricter validation conditions.