Abstract Increased traffic congestion, limited delivery windows, vehicles’ diverse fleets, and dynamic environments are all contributing factors to the inefficiency of urban freight systems. Due to the inability of traditional routing techniques to adjust in real-time to partially visible multi-constrained conditions, delivery delays, fuel consumption, and operating costs all increased. This research proposes a constraint-aware projected policy learning-reinforcement learning (CAPPL-RL) framework to optimize urban freight delivery routes while enforcing real-world constraints, including traffic congestion, delivery time windows, and vehicle capacity. The routing problem is formulated as a partially observable Markov decision process, with autonomous vehicles as agents navigating stochastic urban traffic networks. Using the UFVOD dataset, CAPPL-RL integrates Q-learning with ε-greedy exploration and projection-based constrained policy optimization (PCPO) to enable adaptive, constraint-aware routing. Simulation results demonstrate that CAPPL-RL outperforms PCPO-RL, reducing average delivery time from 65.3 to 52.1 min (20.2%), fuel consumption from 0.12 to 0.093 L/km (22.5%), and constraint violations in time windows from 12 to 3 (75%) while achieving 100% compliance with vehicle capacity limits. These results validate CAPPL-RL as a robust, scalable, and adaptive framework for dynamic urban freight logistics. Similar content being viewed by others Data availability The datasets generated and/or analysed during the current study are available in the [https://www.kaggle.com/datasets/programmer3/urban-freight-and-vehicle-operations-dataset] repository. References Yang, Y., Lu, X., Chen, J. & Li, N. Factor mobility, transportation network, and green economic growth of the urban agglomeration. Sci. Rep. 12(1), 20094. https://doi.org/10.1038/s41598-022-24624-5 (2022). El Yadari, M., Jawab, F., Moufad, I. & Arif, J. Logistics sprawl and urban congestion dynamics toward sustainability: A logistic regression and random-forest-based model. Sustainability 17(13), 5929. https://doi.org/10.3390/su17135929 (2025). Hussain, Q. et al. Reinforcement learning based route optimization model to enhance energy efficiency in internet of vehicles. Sci. Rep. 15(1), 3113. https://doi.org/10.1038/s41598-025-86608-5 (2025). Spyrou, E. D., Kappatos, V., Gkemou, M. & Bekiaris, E. Multimodal transport optimization
Optimization of urban freight <b>intelligent</b> route based on reinforcement learning
Read the original article
nature.com →