Key Takeaways & Executive Findings
- •• The proposed DDPG-based strategy integrates imitation learning and feedforward exploration to enhance learning efficiency and generalization in autonomous vehicle path following. • Imitation learning using MPC-generated data pre-trains the actor network, while feedforward steering with random noise improves exploration. • A hierarchical progressive reward function and a constrained objective reward function inspired by MPC are designed to guide training. • Simulation and HIL tests demonstrate superior path tracking performance and generalization compared to other methods.
Abstract
Autonomous driving technology is constantly developing to a higher level of complex scenes, and there is a growing demand for the utilization of end-to-end data-driven control. However, the end-to-end path tracking process often encounters challenges in learning efficiency and generalization. To address this issue, this paper designs a deep deterministic policy gradient (DDPG)-based reinforcement learning strategy that integrates imitation learning and feedforward exploration in the path following process. In imitation learning, the path tracking control data generated by the model predictive control (MPC) method is used to train an end-to-end steering control model of a deep neural network. Another feedforward exploration behavior is predicted by road curvature and vehicle speed, and adds it and imitation learning to the DDPG reinforcement learning to obtain decision-making experience and action prediction behavior of the path tracking process. In the reinforcement learning process, imitation learning is used to update the pre-training parameters of the actor network, and a feedforward steering technique with random noise is adopted for strategy exploration. In the reward function, a hierarchical progressive reward form and a constrained objective reward function referring to MPC are designed, and the actor-critic network architecture is determined. Finally, the path tracking performance of the designed method is verified by comparing various training results, simulations, and HIL tests. The results show that the designed method can effectively utilize pre-training and feedforward prior experience to obtain optimal path tracking performance of an autonomous vehicle, and has better generalization ability than other methods. This study provides an efficient control scheme for improving the end-to-end control performance of autonomous vehicles.
1. Introduction
Autonomous vehicles are expected to have a significant impact on easing traffic congestion, reducing traffic accidents and freeing human drivers, which has attracted huge research investment from many research institutions and enterprises [1, 2]. In general, the key technologies of autonomous driving mainly focus on environmental perception, decision-making and control execution, among which the path tracking is an essential lateral dynamic control task for autonomous vehicles. High-level autonomous driving technology requires the steering execution of path tracking to work collaboratively with other vehicle subsystems, and designing a path following solution with multi-objective balanced performance will be a necessary prerequisite for highly intelligent autonomous vehicles [3, 4]. Currently, with the increasing level of automation in autonomous vehicles, autonomous driving control issues also create higher demands for motion control accuracy, efficiency, and stability.
In order to comprehensively consider vehicle kinematics and dynamic constraints in performance goals such as vehicle stability and comfort, traditional lateral and longitudinal control methods for autonomous vehicles generally require accurate mathematical analytical models. The commonly used path tracking control strategies include geometric and kinematic control, optimal control, robust control, and model predictive control (MPC) algorithms. These methods require an online solution of the path tracking optimization problem in each control step, which limits the control efficiency to some extent [5–7]. Among the existing methods, pure pursuit and Stanley methods are the most popular geometric controllers to achieve simple and efficient path tracking control behavior. But these types of controllers are generally suitable for smooth path following at low speeds in driving scenarios [8]. To
Loading authentic research manuscript (Pages 1–5)...
Qianjie Liu, Peixiang Xiong, Qingyuan Zhu, Wei Xiao, Kejie Wang, Guoliang Hu, Gang Li (2025). A DDPG-based Path Following Control Strategy for Autonomous Vehicles by Integrated Imitation Learning and Feedforward Exploration. Chinese Journal of Mechanical Engineering. https://doi.org/10.1186/s10033-025-01336-1
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is the main contribution of this paper?
The paper proposes a DDPG-based path following control strategy that integrates imitation learning and feedforward exploration to improve learning efficiency and generalization in autonomous vehicle path tracking.
How does the proposed method improve learning efficiency?
By using imitation learning to pre-train the actor network with MPC-generated data and incorporating feedforward exploration based on road curvature and vehicle speed, the method accelerates learning and enhances exploration.
What are the key components of the reward function?
The reward function includes a hierarchical progressive reward form and a constrained objective reward function inspired by MPC, which helps guide the agent towards optimal path tracking behavior.
How was the proposed method validated?
The method was validated through comparative training results, simulations, and Hardware-in-the-Loop (HIL) tests, demonstrating superior path tracking performance and generalization compared to other methods.
What are the potential applications of this research?
This research provides an efficient control scheme for improving end-to-end control performance of autonomous vehicles, which can be applied in advanced driver assistance systems and fully autonomous driving.
Related Technical Papers & Translations
Direct Repair of the Crystal Structure and Coating Surface of Spent LiFePO4 Materials Enables Superfast Li-Ion Migration
The rapid accumulation of spent LiFePO4 (LFP) cathodes from retired lithium-ion batteries necessitates the development of effective and environmental-friendly recycling strategies. In this context, direct regeneration has emerged as a promising approach for reclaiming LFP cathode materials, offering a streamlined pathway to restore their electrochemical functionality. We report an integrated regeneration protocol that simultaneously repairs the degraded crystal structure and reconstructs the damaged carbon coating in spent LFP. The regenerated cathode material had superfast lithium-ion diffusion kinetics and a stable cathode–electrolyte interface, giving a remarkable rate capability with specific capacities of 122 mAh g−1 at 5C and 106 mAh g−1 at 10C (1C = 170 mA g−1). It also maintained capacities of 110.7 mAh g−1 (5C) and 84.1 mAh g−1 (10C) after 400 cycles. It could be used in harsh environments and could be stably cycled at subzero temperatures (−10 and −20 °C) and in solid-state electrolyte batteries. Life cycle assessment combined with economic evaluation using the EverBatt model reveals that this direct regeneration approach has high economic and environmental benefits.
Oxide Semiconductor for Advanced Memory Architectures: Atomic Layer Deposition, Key Requirement and Challenges
Oxide semiconductors (OSs), introduced by the Hosono group in the early 2000s, have evolved from display backplane materials to promising candidates for advanced memory and logic devices. The exceptionally low leakage current of OSs and compatibility with three-dimensional (3D) architectures have recently sparked renewed interest in their use in semiconductor applications. This review begins by exploring the unique material properties of OSs, which fundamentally originate from their distinct electronic band structure. Subsequently, we focus on atomic layer deposition (ALD), a core technique for growing excellent OS films, covering both basic and advanced processes compatible with 3D scaling. The basic surface reaction mechanisms—adsorption and reaction—and their roles in film growth are introduced. Furthermore, material design strategies, such as cation selection, crystallinity control, anion doping, and heterostructure engineering, are discussed. We also highlight challenges in memory applications, including contact resistance, hydrogen instability, and lack of p-type materials, and discuss the feasibility of ALD-grown OSs as potential solutions. Lastly, we provide an outlook on the role of ALD-grown OSs in memory technologies. This review bridges material fundamentals and device-level requirements, offering a comprehensive perspective on the potential of ALD-driven OSs for next-generation semiconductor memory devices.
Laser powder bed fusion of biodegradable Zn-4Cu alloy: Processing, microstructure and properties
Zn's natural degradability and biocompatibility make it a promising candidate for implants, however, its mechanical properties remain insufficient for bone applications. In this study, the performance of Zn was enhanced by developing Zn-Cu alloys via laser powder bed fusion (LPBF). Optimal LPBF parameters for forming stable tracks were achieved by adjusting laser power and scanning speed. Under optimized conditions of 100 W and 100 mm/s, high-density (99.58%) Zn-Cu alloys with improved hardness (68.2HV) and yield strength (160 MPa) were achieved. These improvements are attributed to solid solution strengthening, segregation strengthening, and grain refinement. The Zn-Cu alloys also demonstrated favorable degradation behavior, with a rate of 0.16 mm/year. This degradation is primarily driven by micro-galvanic corrosion between the CuZn5 phase and Zn matrix, along with refined grains and increased grain boundary density. This work demonstrates a viable strategy for fabricating Zn-based implants with enhanced structural integrity and mechanical performance via LPBF.