计算机科学
强化学习
数学优化
稳健性(进化)
适应性
水准点(测量)
早熟收敛
趋同(经济学)
多目标优化
偏爱
人工智能
机器学习
最优化问题
相似性(几何)
帕累托原理
灵活性(工程)
优化算法
元启发式
局部最优
算法
进化算法
作者
Kaibin Xu,Te Xu,Xianpeng Wang,Erchao Li
标识
DOI:10.1016/j.swevo.2026.102534
摘要
Multiobjective optimization problems (MOPs) require a delicate balance between convergence to the Pareto front (PF) and diversity maintenance. Traditional algorithms often struggle with high-dimensional objectives and a complex PF. This paper introduces a reinforcement learning (RL)-based multi-population co-evolution algorithm with indicator preference guidance (MPCEP), in which each subpopulation plays a specific role. The first subpopulation focuses on improving solution quality through a convergent external archive and a preference evolution strategy, utilizing preference guidance based on solution similarity and reverse evolution to direct solutions towards the true PF. The second subpopulation encourages exploration and ensures diversity by employing a dynamic penalty strategy that prevents premature convergence. The third subpopulation integrates reinforcement learning with a reward–penalty mechanism to promote the emergence of suboptimal solutions, allowing for adaptive preference adjustments during the search process. This co-evolutionary structure facilitates dynamic adjustment for efficient exploration and exploitation of the decision space, while promoting collaboration between subpopulations. Extensive experiments on benchmark problems and real-world problems demonstrate that MPCEP consistently outperforms existing algorithms in terms of both convergence and diversity, highlighting its robustness and adaptability in solving multiobjective optimization problems across various domains.
科研通智能强力驱动
Strongly Powered by AbleSci AI