In the information era, ensuring high-quality data is essential, yet data quality is frequently compromised by human errors, system failures, and other unforeseen factors. Existing tabular anomaly detection methods face several limitations: many are restricted to specific anomaly types or rely heavily on manual labeling, while a few traditional machine learning approaches, despite addressing some of these issues, still encounter substantial challenges due to sample imbalance in practical applications. To address these limitations, we propose MLAD—a fully automated two-stage framework for multi-perspective tabular anomaly detection. In the first stage, preliminary mixed labels are automatically generated from various perspectives to provide essential prior knowledge. In the second stage, a label diffusion module, combined with a multi-instance learning strategy, effectively alleviates the impact of data imbalance and enables comprehensive anomaly detection from multiple views. Experimental results on six benchmark datasets demonstrate that MLAD achieves high detection accuracy and automation without requiring human intervention or external database support.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MLAD: A Two-Stage Fully Automated Multi-perspective Tabular Data Anomaly Detection Framework

  • Hongtao Song,
  • Yufei Li,
  • Qilong Han

摘要

In the information era, ensuring high-quality data is essential, yet data quality is frequently compromised by human errors, system failures, and other unforeseen factors. Existing tabular anomaly detection methods face several limitations: many are restricted to specific anomaly types or rely heavily on manual labeling, while a few traditional machine learning approaches, despite addressing some of these issues, still encounter substantial challenges due to sample imbalance in practical applications. To address these limitations, we propose MLAD—a fully automated two-stage framework for multi-perspective tabular anomaly detection. In the first stage, preliminary mixed labels are automatically generated from various perspectives to provide essential prior knowledge. In the second stage, a label diffusion module, combined with a multi-instance learning strategy, effectively alleviates the impact of data imbalance and enables comprehensive anomaly detection from multiple views. Experimental results on six benchmark datasets demonstrate that MLAD achieves high detection accuracy and automation without requiring human intervention or external database support.