Journal of Systems Engineering and Electronics ›› 2026, Vol. 37 ›› Issue (4): 1297-1316.doi: 10.23919/JSEE.2026.000009

• SYSTEMS ENGINEERING • Previous Articles    

Curriculum learning for adaptive multi-UAV swarm target search in obstacle environment

Xin Cao1,2(), He Luo1,2,*(), Guoqiang Wang1,2(), Yingying Ma1,2(), Jue Zhang3()   

  1. 1School of Management, Hefei University of Technology, Hefei 230009, China
    2Key Laboratory of Process Optimization Intelligent Decision-making, Ministry of Education, Hefei 230009, China
    3ARC Research Hub for Driving Farming Productivity and Disease Prevention, Griffith University, Queensland 4111, Australia
  • Received:2025-05-27 Online:2026-08-18 Published:2026-09-03
  • Contact: He Luo E-mail:caoxin@mail.hfut.edu.cn;luohe@hfut.edu.cn;gqwang2017@hfut.edu.cn;mayingying@mail.hfut.edu.cn;jue.zhang@griffith.edu.au
  • Supported by:
    This work was supported by the National Natural Science Foundation of China (72571094; 71971075; 72271076; 71871079; 71671059).

Abstract:

Previous studies have paid limited attention to the critical role of unmanned aerial vehicle (UAV) altitude variation in determining three-dimensional (3D) search efficiency. To address this gap, this paper introduces a deep reinforcement learning framework integrated with adaptive matrix optimization in multi-UAV target search. The proposed framework explicitply models altitude variations’ impacts on observation range and detection accuracy. To balance flight safety, coverage, and detection precision, deep Q-network (DQN)-altitude matrix optimization (AMO) discretizes 3D trajectories into horizontal paths and vertical altitude decisions, effectively reducing the problem dimensionality. Furthermore, a curriculum learning approach is employed to decompose the search process into phased sub-tasks, each with customized decision rules and reward mechanisms. This hierarchical strategy accelerates agent learning and enhances performance in complex scenarios. Comprehensive experimental evaluations in simulated environments demonstrate that the proposed DQN-AMO outperforms benchmark methods in both robustness and generalization.

Key words: multi-UAV swarm, adaptive optimization, target search, curriculum learning