• The proposed DDPG-based strategy integrates imitation learning and feedforward exploration to enhance learning efficiency and generalization in autonomous vehicle path following.
• Imitation learning using MPC-generated data pre-trains the actor network, while feedforward steering with random noise improves exploration.
• A hierarchical progressive reward function and a constrained objective reward function inspired by MPC are designed to guide training.
• Simulation and HIL tests demonstrate superior path tracking performance and generalization compared to other methods.