•  
  •  
 

Journal of System Simulation

Abstract

Abstract: To address the problems such as low exploration and sampling efficiency and sparse rewards in the early training stage under the complex target pursuit environment of multiple AUVs based on MADDPG, a phased-guidance curriculum MADDPG (PGC-MADDPG) algorithm was proposed. The pursuit task was divided into two phases, i.e., target tracking and encircling, through curriculum learning. In the target tracking phase, an experience strategy based on the APF method was introduced as a guidance item to provide prior knowledge of the target direction and accelerate the AUV training speed. After entering the encircling phase, the APF guidance item was removed, enabling the AUVs to accomplish the pursuit task relying on their own strategies. Comparative experimental results indicate that compared with the curriculum learning MADDPG algorithm, the proposed algorithm improves the convergence speed by approximately 36% and increases the final pursuit success rate by approximately 5%.

First Page

1964

Last Page

1977

CLC

TP273

Recommended Citation

Zhang Sen, Shen Sihang, Sun Xiaojie, et al. Multi-AUV Pursuit Algorithm with Phased Guidance Based on MADDPG[J]. Journal of System Simulation, 2026, 38(7): 1964-1977.

Corresponding Author

Sun Xiaojie

DOI

10.16182/j.issn1004731x.joss.25-0894

Share

COinS