Adaptive causal discovery with ordering–based reinforcement learning-approach for type 2 diabetes mellitus risk factors
摘要
Causal discovery of type 2 diabetes mellitus (T2DM) risk factors is crucial for revealing underlying disease mechanisms and supporting prevention and treatment. However, clinical risk-factor data are typically high-dimensional, limiting the performance of many conventional causal discovery methods. Causal discovery with ordering-based reinforcement learning (CORL) reduces the search space via variable orderings, but it uniformly treats all variable states during initial state construction, failing to account for the heterogeneous importance of individual risk factors.
MethodsTo address these issues, this study proposes an adaptive CORL framework (ACORL) tailored for T2DM risk factors. First, the search for a causal graph is reformulated into a variable ordering task, where feature contributions are estimated via an ensemble Random Forest classifier to construct a weighted initial state. Next, the variable ordering is optimized through a reinforcement learning reward mechanism to generate a directed acyclic graph. Subsequently, ordinary least squares linear regression pruning and empirical information entropy estimation are applied to eliminate redundant relations and quantify causal strengths. Finally, comparative experiments were conducted on both synthetic and empirical clinical datasets to evaluate structural accuracy and medical plausibility.
ResultsSimulations on a synthetic dataset demonstrated that ACORL reduced the structural Hamming distance from 3 to 1 and enhanced the F1 score from 50% to 85.71% relative to CORL. On empirical clinical data, ACORL identified causal relations highly consistent with established medical knowledge, achieving a correct-edge rate of 83.3% on the 13-feature NHANES dataset and 91.6% on the 29-feature NHANES dataset. Furthermore, it matched the baseline performance on the lower-dimensional Pima datasets while requiring significantly fewer training iterations.
ConclusionsThis study introduces an adaptive reinforcement learning framework for T2DM causal discovery. The proposed ACORL model effectively leverages feature heterogeneity and statistical refinement to discover medically plausible causal relations among risk factors. Consequently, it provides an efficient and reliable tool for epidemiological researchers investigating risk-factor structures in large-scale clinical studies.