<p>Offline constraint-based causal feature selection (OC-CFS) algorithms are essential for identifying causal relationships from observational data. However, existing methods often suffer from limitations such as low prediction accuracy or high computational cost, particularly when sample sizes vary. To address these limitations, we propose Triplet, a novel framework that leverages the HITON-MB Parents and Children (PC) strategy to identify strongly relevant PC nodes while eliminating irrelevant and redundant features. It concurrently employs the BAMB strategy to detect relevant spouses and discard irrelevant ones, and applies the STMB non-Markov Blanket (non-MB) strategy to identify and exclude non-MB descendants. Through this integration, the proposed T-OCD<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_24421_Article_IEq1.gif" Format="GIF" Height="11" Rendition="HTML" Resolution="72" Type="Linedraw" Width="25" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{MB}\)</EquationSource> </InlineEquation> overcomes these limitations, accurately identifying the true MB with high prediction accuracy and reduced runtime. To validate its effectiveness, we evaluated T-OCD<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_24421_Article_IEq1.gif" Format="GIF" Height="11" Rendition="HTML" Resolution="72" Type="Linedraw" Width="25" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{MB}\)</EquationSource> </InlineEquation> on benchmark Bayesian networks (BNs) and real-world datasets. Extensive experimental results demonstrate that T-OCD<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_24421_Article_IEq1.gif" Format="GIF" Height="11" Rendition="HTML" Resolution="72" Type="Linedraw" Width="25" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{MB}\)</EquationSource> </InlineEquation> achieves significant improvements in both prediction accuracy and computational efficiency compared to existing methods. On small sample sizes (n=500), T-OCD<InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_24421_Article_IEq1.gif" Format="GIF" Height="11" Rendition="HTML" Resolution="72" Type="Linedraw" Width="25" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{MB}\)</EquationSource> </InlineEquation> achieved the highest recall in 5 out of 7 datasets, with an average improvement of over 20% compared to rivals. On large sample sizes (n=5000), it excelled in precision, achieving the top score in 4 out of 7 datasets with an average precision of 94%. Computationally, T-OCD<InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_24421_Article_IEq1.gif" Format="GIF" Height="11" Rendition="HTML" Resolution="72" Type="Linedraw" Width="25" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{MB}\)</EquationSource> </InlineEquation> is highly efficient, operating as the second-fastest method overall. It ran over 55% faster than half of the benchmarks and a remarkable 35% faster than the average competitor on large datasets. The source code for this research is available at the following repository: <a href="https://github.com/vickykhan89/T-OCDmb.">https://github.com/vickykhan89/T-OCDmb.</a></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Triplet offline causal discovery based on optimal Markov blanket and its application

  • Waqar Khan,
  • Brekhna Brekhna,
  • Jianqiong Huang,
  • Yajun Xie,
  • Muhammad Suhail Shaikh,
  • Muhammad Sadiq Hassan Zada,
  • Yifan Zheng

摘要

Offline constraint-based causal feature selection (OC-CFS) algorithms are essential for identifying causal relationships from observational data. However, existing methods often suffer from limitations such as low prediction accuracy or high computational cost, particularly when sample sizes vary. To address these limitations, we propose Triplet, a novel framework that leverages the HITON-MB Parents and Children (PC) strategy to identify strongly relevant PC nodes while eliminating irrelevant and redundant features. It concurrently employs the BAMB strategy to detect relevant spouses and discard irrelevant ones, and applies the STMB non-Markov Blanket (non-MB) strategy to identify and exclude non-MB descendants. Through this integration, the proposed T-OCD \(_{MB}\) overcomes these limitations, accurately identifying the true MB with high prediction accuracy and reduced runtime. To validate its effectiveness, we evaluated T-OCD \(_{MB}\) on benchmark Bayesian networks (BNs) and real-world datasets. Extensive experimental results demonstrate that T-OCD \(_{MB}\) achieves significant improvements in both prediction accuracy and computational efficiency compared to existing methods. On small sample sizes (n=500), T-OCD \(_{MB}\) achieved the highest recall in 5 out of 7 datasets, with an average improvement of over 20% compared to rivals. On large sample sizes (n=5000), it excelled in precision, achieving the top score in 4 out of 7 datasets with an average precision of 94%. Computationally, T-OCD \(_{MB}\) is highly efficient, operating as the second-fastest method overall. It ran over 55% faster than half of the benchmarks and a remarkable 35% faster than the average competitor on large datasets. The source code for this research is available at the following repository: https://github.com/vickykhan89/T-OCDmb.