openHSU logo
Log In(current)
  1. Home
  2. Helmut-Schmidt-University / University of the Federal Armed Forces Hamburg
  3. Publications
  4. 3 - Publication references (without full text)
  5. Using deep reinforcement learning with automatic curriculum learning for mapless navigation in intralogistics

Using deep reinforcement learning with automatic curriculum learning for mapless navigation in intralogistics

Publication date
2022-03-19
Document type
Forschungsartikel
Author
Xue, Honghu
Hein, Benedikt
Bakr, Mohamed
Schildbach, Georg
Abel, Bengt
Rueckert, Elmar
Organisational unit
Technologie von Logistiksystemen  
DOI
10.3390/app12063153
URI
https://openhsu.ub.hsu-hh.de/handle/10.24405/17886
Publisher
MDPI
Series or journal
Applied Sciences
ISSN
2076-3417
Periodical volume
12
Periodical issue
6
Article ID
3153
Peer-reviewed
✅
Part of the university bibliography
✅
Additional Information
Language
English
Keyword
Deep reinforcement learning
Automatic curriculum learning
Autonomous navigation
Multi-modal sensor perception
Abstract
We propose a deep reinforcement learning approach for solving a mapless navigation problem in warehouse scenarios. In our approach, an automatic guided vehicle is equipped with two LiDAR sensors and one frontal RGB camera and learns to perform a targeted navigation task. The challenges reside in the sparseness of positive samples for learning, multi-modal sensor perception with partial observability, the demand for accurate steering maneuvers together with long training cycles. To address these points, we propose NavACL-Q as an automatic curriculum learning method in combination with a distributed version of the soft actor-critic algorithm. The performance of the learning algorithm is evaluated exhaustively in a different warehouse environment to validate both robustness and generalizability of the learned policy. Results in NVIDIA Isaac Sim demonstrates that our trained agent significantly outperforms the map-based navigation pipeline provided by NVIDIA Isaac Sim with an increased agent-goal distance of 3 m and a wider initial relative agent-goal rotation of approximately 45∘. The ablation studies also suggest that NavACL-Q greatly facilitates the whole learning process with a performance gain of roughly 40% compared to training with random starts and a pre-trained feature extractor manifestly boosts the performance by approximately 60%.
Description
This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).
Version
Published version
Access right on openHSU
Metadata only access

  • Privacy policy
  • Send Feedback
  • Imprint