The classification of dispersed data poses challenges due to inconsistencies and conflicts arising from independently collected sources. This study introduces a coalition-based classification framework that integrates conflict analysis and rule-based learning. The approach employs four decision rule induction methods – exhaustive search, genetic algorithms, covering algorithms, and LEM2 – combined with three decision-making strategies: first rule approach, all rules approach, and weighted rules approach. Experiments were conducted on datasets from the UCI Machine Learning Repository. The theoretical contribution of the paper is a novel classification structure for dispersed data, which utilizes conflict analysis to identify consistent sources and form coalitions. The practical contribution involves the development of an interpretable method that enables the generation of transparent rules and allows comparison of different approaches in the context of dispersed data. Results indicate that the covering algorithm with weighted rules approach achieves the highest classification performance across all metrics. The limitations of the study include the poor performance of the LEM2 method, which often fails to generate covering rules, leading to random classifications, as well as aspects of scalability that may need further attention for large datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating Conflict Analysis and Rule-Based Systems for Dispersed Data Classification

  • Małgorzata Przybyła-Kasperek,
  • Katarzyna Kusztal

摘要

The classification of dispersed data poses challenges due to inconsistencies and conflicts arising from independently collected sources. This study introduces a coalition-based classification framework that integrates conflict analysis and rule-based learning. The approach employs four decision rule induction methods – exhaustive search, genetic algorithms, covering algorithms, and LEM2 – combined with three decision-making strategies: first rule approach, all rules approach, and weighted rules approach. Experiments were conducted on datasets from the UCI Machine Learning Repository. The theoretical contribution of the paper is a novel classification structure for dispersed data, which utilizes conflict analysis to identify consistent sources and form coalitions. The practical contribution involves the development of an interpretable method that enables the generation of transparent rules and allows comparison of different approaches in the context of dispersed data. Results indicate that the covering algorithm with weighted rules approach achieves the highest classification performance across all metrics. The limitations of the study include the poor performance of the LEM2 method, which often fails to generate covering rules, leading to random classifications, as well as aspects of scalability that may need further attention for large datasets.