HRL-MOEA: a hybrid reinforcement learning-enhanced multi-objective recommendation algorithm with dynamic policy orchestration
摘要
Recommendation algorithms have become increasingly prevalent in modern society, addressing information overload by delivering content aligned with user preferences. While traditional approaches prioritize recommendation accuracy, singular focus on this objective often results in popularity bias. This imbalance introduces fairness concerns for item providers and detrimental feedback loops in recommendation ecosystems, highlighting the critical importance of item exposure fairness. However, balancing these dual objectives faces fundamental trade-off challenges. Existing multi-objective recommendation methods typically rely on empirically fixed genetic operators during evolutionary processes, which not only requires laborious parameter tuning but also constrains the generation of high-quality solutions. To overcome these limitations, we propose a hybrid reinforcement learning-enhanced adaptive evolutionary algorithm (HRL-MOEA). The framework synergistically integrates SARSA and Q-learning strategies through a phase-aware mechanism: during early evolutionary stages, a conservative SARSA-based self-adaptive mechanism facilitates comprehensive solution space exploration, while strategically transitioning to Q-learning’s exploitation-oriented policy in later phases to accelerate convergence. This dynamic strategy is conducive to enhancing the model’s evolutionary performance. Experimental results demonstrate that HRL-MOEA outperforms existing algorithms in performance effectiveness.