Machine Learning · 2018 to 2019
Deep Reinforcement Learning for Autonomous Vehicles
Master's thesis training PPO agents to drive in the CARLA simulator, with a focus on cutting training time.




1 / 4
Eager to explore a technology that promises to change society, I joined the autonomous vehicle lab at NTNU for my final two semesters. Together with my supervisor, I decided to explore reinforcement learning for autonomous vehicles.
The prospect of an AI that learns to drive and play games by trial and error fascinated me, and the project let me combine my experience with game engines and computer vision. Reinforcement learning for autonomous vehicles typically relies on driving simulators such as CARLA or AirSim, both of which run on Unreal Engine 4.
For the practical part of the project, I implemented OpenAI’s Proximal Policy Optimization algorithm, which in 2018 was the baseline for general-purpose reinforcement learning.
The main finding of my preliminary study (also see this video) was that scaling the means of the Gaussian action distributions to the range of valid actions had a huge impact. For example, if action 0 controls steering, its valid range might be [-45, 45] degrees. If the network outputs an unbounded mean (effectively [-∞, ∞]), it takes much longer to converge. Scaling the means to the appropriate range for every action substantially increased training speed. To the best of my knowledge, the authors of PPO do not do this, and it is not in OpenAI’s official PPO code.
The initial test environment was admittedly simplistic, so I built a custom RL environment in CARLA, which is public in this repository. The more complex environment drastically increased training time, and I spent the second half of my master’s searching for a model that would learn to drive reliably within a day. This video shows the results of those experiments, and the final report goes into detail.