Implementation of a reinforcement learning-based approach for autonomous and efficient locomotion of wheeled mobile robots
DOI:
https://doi.org/10.70929/caui3.v2i1.0015Keywords:
Reinforcement learning, PPO, turtlebot, simulation, wheeled robotAbstract
This study aims to optimize the autonomous navigation of a Turtlebot wheeled mobile robot in a 2D maze environment. The main objective focuses on minimizing both the number of steps and the time required to reach a target, for which the benefits of a reinforcement learning algorithm have been exploited. Specifically, the Proximal Policy Optimization (PPO) learning algorithm has been selected. Thus, the robot demonstrated an exponential learning curve, achieving a high level of effectiveness in a short period of time. A continuous and significant improvement in the quality of the trajectories generated has been observed, as well as a progressive reduction in the number of movements required to complete the task throughout the training sessions. This training process has been organized into consecutive sessions, consisting of ten thousand movements, with the ability to save the learned model for further refinement in subsequent iterations. The results validate that the PPO algorithm is a highly promising methodology for autonomous motion control in robots.
References
[1] M. Häselich, M. Arends, N. Wojke, F. Neuhaus, and D. Paulus, “Probabilistic terrain classification in unstructured environments,” Rob. Auton. Syst., vol. 61, no. 10, pp. 1051–1059, Oct. 2013, doi: 10.1016/j.robot.2012.08.002.
[2] L. Wijayathunga, A. Rassau, and D. Chai, “Challenges and Solutions for Autonomous Ground Robot Scene Understanding and Navigation in Unstructured Outdoor Environments: A Review,” Applied Sciences, vol. 13, no. 17, p. 9877, Aug. 2023, doi: 10.3390/app13179877.
[3] F. Qi, “Local probabilistic terrain estimation for navigation in unstructured environments,” in Third International Conference on Machine Vision, Automatic Identification, and Detection (MVAID 2024), R. Jin, Ed., SPIE, Aug. 2024, p. 120. doi: 10.1117/12.3036525.
[4] M. J. Procopio, J. Mulligan, and G. Grudic, “Learning terrain segmentation with classifier ensembles for autonomous robot navigation in unstructured environments,” J. Field Robot., vol. 26, no. 2, pp. 145–175, Feb. 2009, doi: 10.1002/rob.20279.
[5] A. Cantos-Andrade and F. Samaniego, “Diseño y construcción de chasis para un robot móvil de ruedas para la exploración de terrenos irregulares,” Revista Digital Científica Causalidad (caui3), vol. 1, no. 1, pp. 1–1, 2025, doi: 10.70929/caui3.v1i1.2wrkyz89.
[6] B. K. Patle, G. Babu L, A. Pandey, D. R. K. Parhi, and A. Jagadeesh, “A review: On path planning strategies for navigation of mobile robot,” Defence Technology, vol. 15, no. 4, pp. 582–606, Aug. 2019, doi: 10.1016/j.dt.2019.04.011.
[7] A. N. A. Rafai, N. Adzhar, and N. I. Jaini, “A Review on Path Planning and Obstacle Avoidance Algorithms for Autonomous Mobile Robots,” Journal of Robotics, vol. 2022, pp. 1–14, Dec. 2022, doi: 10.1155/2022/2538220.
[8] L. G. Galvao, M. Abbod, T. Kalganova, V. Palade, and M. N. Huda, “Pedestrian and Vehicle Detection in Autonomous Vehicle Perception Systems—A Review,” Sensors, vol. 21, no. 21, p. 7267, Oct. 2021, doi: 10.3390/s21217267.
[9] J. Crespo, J. C. Castillo, O. M. Mozos, and R. Barber, “Semantic Information for Robot Navigation: A Survey,” Applied Sciences, vol. 10, no. 2, p. 497, Jan. 2020, doi: 10.3390/app10020497.
[10] S. Levine, P. Pastor, A. Krizhevsky, J. Ibarz, and D. Quillen, “Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection,” Int. J. Rob. Res., vol. 37, no. 4–5, pp. 421–436, Apr. 2018, doi: 10.1177/0278364917710318.
[11] S. Benford, C. Magerkurth, and P. Ljungstrand, “Bridging the physical and digital in pervasive gaming,” Commun. ACM, vol. 48, no. 3, pp. 54–57, Mar. 2005, doi: 10.1145/1047671.1047704.
[12] M. N. Ahangar, Q. Z. Ahmed, F. A. Khan, and M. Hafeez, “A Survey of Autonomous Vehicles: Enabling Communication Technologies and Challenges,” Sensors, vol. 21, no. 3, p. 706, Jan. 2021, doi: 10.3390/s21030706.
[13] A. Sarker et al., “A Review of Sensing and Communication, Human Factors, and Controller Aspects for Information-Aware Connected and Automated Vehicles,” IEEE Transactions on Intelligent Transportation Systems, vol. 21, no. 1, pp. 7–29, Jan. 2020, doi: 10.1109/TITS.2019.2892399.
[14] Ł. Łach and D. Svyetlichnyy, “Comprehensive Review of Traffic Modeling: Towards Autonomous Vehicles,” Applied Sciences, vol. 14, no. 18, p. 8456, Sep. 2024, doi: 10.3390/app14188456.
[15] W. Qiang, L. Yang, and H. Jin, “Efficient and Robust Malware Detection Based on Control Flow Traces Using Deep Neural Networks,” Comput. Secur., vol. 122, p. 102871, Nov. 2022, doi: 10.1016/j.cose.2022.102871.
[16] Y. Gu, Y. Cheng, C. L. P. Chen, and X. Wang, “Proximal Policy Optimization With Policy Feedback,” IEEE Trans. Syst. Man Cybern. Syst., vol. 52, no. 7, pp. 4600–4610, Jul. 2022, doi: 10.1109/TSMC.2021.3098451.
[17] N. Li et al., “A Progress Review on Solid‐State LiDAR and Nanophotonics‐Based LiDAR Sensors,” Laser Photon. Rev., vol. 16, no. 11, Nov. 2022, doi: 10.1002/lpor.202100511.
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2026 Scientific Digital Journal Causality Causality (caui3)

This work is licensed under a Creative Commons Attribution 4.0 International License.







