1.广东轻工职业技术大学智能制造与装备学院,广东 广州 510399
2.中山大学微电子科学与技术学院, 广东 珠海 519082
3.中山大学智能工程学院, 广东 深圳 518107
陈永灿(1988年生),男;研究方向:深度强化学习;E-mail:chenycan@alumni.sysu.edu.cn
唐承佩(1978年生),男;研究方向:边缘智能;E-mail:tchengp@mail.sysu.edu.cn
收稿:2026-04-30,
录用:2026-06-11,
网络首发:2026-09-21,
纸质出版:2026-09-25
移动端阅览
陈永灿,黄晓,梁海澄等.基于改进型DDPG的AGV自主智能导航算法[J].中山大学学报(自然科学版)(中英文),2026,65(05):118-128.
Chen Yongcan,Huang Xiao,Liang Haicheng,et al.AGV autonomous intelligent navigation algorithm based on improved DDPG[J].Acta Scientiarum Naturalium Universitatis Sunyatseni,2026,65(05):118-128.
陈永灿,黄晓,梁海澄等.基于改进型DDPG的AGV自主智能导航算法[J].中山大学学报(自然科学版)(中英文),2026,65(05):118-128. DOI: 10.11714/acta.snus.ZR20260118.
Chen Yongcan,Huang Xiao,Liang Haicheng,et al.AGV autonomous intelligent navigation algorithm based on improved DDPG[J].Acta Scientiarum Naturalium Universitatis Sunyatseni,2026,65(05):118-128. DOI: 10.11714/acta.snus.ZR20260118.
提出了基于改进型深度确定性策略梯度(DDPG)的自动导引车(AGV)自主智能导航算法。首先,提出了基于环境风险的自适应探索策略,其中包含对动态障碍物轨迹的预判;其次,设计了多目标加权的奖励函数,指引AGV在避开障碍物且不越界的情况下尽快到达目的地;再次,基于风险系数-经验新鲜度双因素抽取经验,缩短训练周期,提升训练质量;最后,运行时的动作预演进一步降低碰撞和越界风险。为验证算法的性能,在不同障碍物数量、尺寸、移动速度条件下开展仿真实验。实验结果显示,基于改进型DDPG的AGV导航算法平均成功率达94.8%,平均超时率低于3%,平均碰撞率低于2.3%;训练周期和导航耗时亦短于对比算法,具有明显优势。
An autonomous intelligent navigation algorithm for automatic guided vehicles(AGV) based on an improved deep deterministic policy gradient(DDPG) is proposed.First,an adaptive exploration strategy based on environmental risk is introduced,which includes predicting the trajectories of dynamic obstacles. Next, a multi-objective weighted reward function is designed to guide the AGV to reach its destination as quickly as possible while avoiding obstacles and staying within boundaries.Then,experiences are extracted using a dual-factor approach based on risk coefficients and experience freshness,which shortens the training cycle and improves training quality. Finally,action rehearsal during runtime further reduces the risk of collisions and boundary violations. To validate the algorithm's performance,simulation experiments were conducted under different numbers,sizes,and speeds of obstacles.The results showed that the AGV navigation algorithm based on the improved DDPG achieved an average success rate of 94.8%, an average timeout rate below 3%,and an average collision rate below 2.3%. Moreover,both the training cycles and navigation time were shorter than those of comparison algorithms, showing clear advantages.
王贺 , 许佳宁 , 闫广宇 , 2025 . 基于深度强化学习的AGV行人避让策略研究 [J]. 系统仿真学报 , 37 ( 3 ): 595 - 606 .
王唯鉴 , 2023 . AGV动态避障及任务调度深度强化学习研究 [D]. 北京 : 机械科学研究总院 .
Arulkumaran K , Deisenroth M P , Brundage M , et al , 2017 . Deep reinforcement learning: A brief survey [J]. IEEE Signal Process Mag , 34 ( 6 ): 26 - 38 .
Chen X Q , Liu S H , Li C F , et al , 2023 . AGV path planning and optimization with deep reinforcement learning model [C]// 7th International Conference on Transportation Information and Safety . Xi'an, China : 1859 - 1863 .
Gao T H , Chen B C , Mi Q W , 2022 . A survey of Markov model in reinforcement learning [C]// 2022 International Conference on Artificial Intelligence in Information and Communication . Jeju Island, Republic of Korea : 284 - 287 .
Hafiz A M , 2022 . A survey of deep Q-networks used for reinforcement learning:State of the art [C]// Intelligent Communication Technologies and Virtual Mobile Networks:Proceedings of ICICV 2022 . Singapore : 393 - 402 .
Lillicrap T P , Hunt J J , Pritzel A , et al , 2015 . Continuous control with deep reinforcement learning [PP/OL].[ 2026-04-03 ]. https://doi.org/10.48550/arXiv.1509.02971 https://doi.org/10.48550/arXiv.1509.02971 .
Liu N , Ma C Y , Hu Z H , et al , 2024 . Workshop AGV path planning based on improved A * algorithm [J]. Math Biosci Eng , 21 ( 2 ): 2137 - 2162 .
Shen G C , Ma R , Tang Z , et al , 2021 . A deep reinforcement learning algorithm for warehousing multi-AGV path planning [C]// International Conference on Networking,Communications and Information Technology . Manchester, United Kingdom : 421 - 429 .
Sutton R S , Barto A G , 2018 . Reinforcement learning: An introduction [M]. Cambridge : The MIT Press .
Wang X , Wang S , Liang X X , et al , 2024 . Deep reinforcement learning: A survey [J]. IEEE Trans Neural Netw Learn Syst , 35 ( 4 ): 5064 - 5078 .
Wu H , 2025 . Research on AGV path planning algorithm integrating adaptive A ∗ and improved APF algorithm [C]// 8th International Conference on Advanced Algorithms and Control Engineering . Shanghai, China : 764 - 769 .
Zhang B , Zhu M W , Lin C H , et al , 2022 . Research on AGV map building and positioning based on SLAM technology [C]// 5th International Conference on Automation, Electronics and Electrical Engineering . Shenyang,China : 707 - 713 .
0
浏览量
10
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621
