Adaptive Successor Features for Transfer Reinforcement Learning

Publication Type:
Thesis
Issue Date:
2025
Full metadata record
Reinforcement Learning has demonstrated remarkable results as a paradigm for machine learning, learning from trial-and-error interactions with an environment. Nevertheless, this paradigm requires an extensive number of samples to perform a specific task, and for a new, unseen task the learning process must start from scratch. Transfer Reinforcement Learning (TRL) is an emerging paradigm that aims to improve sample efficiency by reusing and adapting previously learnt knowledge in new tasks. In recent years, Successor Features (SFs) have been a widely studied framework for TRL. SFs are a robust mechanism for transfer learning that excel in scenarios where rewards are a linear combination of features. They achieve transfer by decoupling transition dynamics from rewards. Moreover, together with Generalised policy improvement (GPI), they form a powerful framework able to compose and transfer solutions to new tasks. This framework has shown promising results in diverse applications such as neuroscience, robotics, autonomous driving, and control. Despite these outstanding results and applications, SFs and GPI present two main limitations. First, they assume that transition dynamics are fixed and shared across tasks, an assumption that hinders knowledge transfer in realistic scenarios where subtle variations in physical or environmental factors can lead to significantly different dynamics. Second, the performance of SFs and GPI on target tasks is theoretically tied to their distance from the source tasks, which limits their ability to guarantee optimal policies for more distant tasks. These challenges highlight the need for new composition mechanisms. This thesis explores mechanisms to enhance SFs and GPI, addressing their current limitations and making the framework more suitable for complex real-world applications. Furthermore, it considers a novel setting in which Reinforcement Learning (RL) is not feasible, showing how SFs can make RL applicable to support human learning.
Please use this identifier to cite or link to this item: