<p>A rigorous game-theoretic framework for modeling stealthy Advanced Persistent Threat (APT) propagation and adaptive defensive response in heterogeneous 6G-enabled edge environments is presented. The attacker–defender interaction is formulated as an Asymmetric Partially Observable Stochastic Game (POSG), capturing divergent observation capabilities, state-dependent action feasibility, and incomplete information. The joint state space is factored to encode 6G architectural heterogeneity (UE, MEC, gNB, CoreNF) alongside multi-stage APT kill-chain semantics, network slice compartmentalization, THz channel dynamics, and persistent telemetry variables. Policies are optimized toward an ε-Asymmetric Perfect Bayesian Equilibrium (ε-PBE) via deep reinforcement learning, with a GNN + GRU Neural Belief Encoder prescribed for scalable deployment. Simulations reveal emergent strategic belief deception and non-monotone learning trajectories consistent with ε-PBE dynamics. While the full model prioritizes long-term strategic robustness over short-term reward maximization (mean defender reward ≈ − 450), it achieves a + 33.5% reward improvement and − 24.7% exfiltration reduction over heuristic baselines. Ablation analysis confirms that target-aware mitigation, compromise depth progression, and Harsanyi type-belief updating yield non-redundant strategic benefits, underscoring that no single metric captures the full complexity of asymmetric 6G-edge defense.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Asymmetric partially observable Markov decision process (POMDP)-based stochastic games for modeling stealthy APT propagation in heterogeneous 6G-enabled edge environments

  • Jude Njoku

摘要

A rigorous game-theoretic framework for modeling stealthy Advanced Persistent Threat (APT) propagation and adaptive defensive response in heterogeneous 6G-enabled edge environments is presented. The attacker–defender interaction is formulated as an Asymmetric Partially Observable Stochastic Game (POSG), capturing divergent observation capabilities, state-dependent action feasibility, and incomplete information. The joint state space is factored to encode 6G architectural heterogeneity (UE, MEC, gNB, CoreNF) alongside multi-stage APT kill-chain semantics, network slice compartmentalization, THz channel dynamics, and persistent telemetry variables. Policies are optimized toward an ε-Asymmetric Perfect Bayesian Equilibrium (ε-PBE) via deep reinforcement learning, with a GNN + GRU Neural Belief Encoder prescribed for scalable deployment. Simulations reveal emergent strategic belief deception and non-monotone learning trajectories consistent with ε-PBE dynamics. While the full model prioritizes long-term strategic robustness over short-term reward maximization (mean defender reward ≈ − 450), it achieves a + 33.5% reward improvement and − 24.7% exfiltration reduction over heuristic baselines. Ablation analysis confirms that target-aware mitigation, compromise depth progression, and Harsanyi type-belief updating yield non-redundant strategic benefits, underscoring that no single metric captures the full complexity of asymmetric 6G-edge defense.