Journal of System Simulation
Abstract
Abstract: To address the problems such as low exploration and sampling efficiency and sparse rewards in the early training stage under the complex target pursuit environment of multiple AUVs based on MADDPG, a phased-guidance curriculum MADDPG (PGC-MADDPG) algorithm was proposed. The pursuit task was divided into two phases, i.e., target tracking and encircling, through curriculum learning. In the target tracking phase, an experience strategy based on the APF method was introduced as a guidance item to provide prior knowledge of the target direction and accelerate the AUV training speed. After entering the encircling phase, the APF guidance item was removed, enabling the AUVs to accomplish the pursuit task relying on their own strategies. Comparative experimental results indicate that compared with the curriculum learning MADDPG algorithm, the proposed algorithm improves the convergence speed by approximately 36% and increases the final pursuit success rate by approximately 5%.
Recommended Citation
Zhang, Sen; Shen, Sihang; Sun, Xiaojie; Shao, Jingping; Guo, Shuaiqiang; and Deng, Yingjie
(2026)
"Multi-AUV Pursuit Algorithm with Phased Guidance Based on MADDPG,"
Journal of System Simulation: Vol. 38:
Iss.
7, Article 12.
DOI: 10.16182/j.issn1004731x.joss.25-0894
Available at:
https://dc-china-simulation.researchcommons.org/journal/vol38/iss7/12
First Page
1964
Last Page
1977
CLC
TP273
Recommended Citation
Zhang Sen, Shen Sihang, Sun Xiaojie, et al. Multi-AUV Pursuit Algorithm with Phased Guidance Based on MADDPG[J]. Journal of System Simulation, 2026, 38(7): 1964-1977.
DOI
10.16182/j.issn1004731x.joss.25-0894
Included in
Artificial Intelligence and Robotics Commons, Computer Engineering Commons, Numerical Analysis and Scientific Computing Commons, Operations Research, Systems Engineering and Industrial Engineering Commons, Systems Science Commons