Feiden, Arno Berengar: Diversity in Reinforcement Learning. - Bonn, 2026. - Dissertation, Rheinische Friedrich-Wilhelms-Universität Bonn.
Online-Ausgabe in bonndoc: https://nbn-resolving.org/urn:nbn:de:hbz:5-91296
Online-Ausgabe in bonndoc: https://nbn-resolving.org/urn:nbn:de:hbz:5-91296
@phdthesis{handle:20.500.11811/14349,
urn: https://nbn-resolving.org/urn:nbn:de:hbz:5-91296,
author = {{Arno Berengar Feiden}},
title = {Diversity in Reinforcement Learning},
school = {Rheinische Friedrich-Wilhelms-Universität Bonn},
year = 2026,
month = aug,
note = {Reinforcement learning is a machine learning paradigm that aims to develop optimal policies for discrete-time control problems based on quantitative feedback on the learner’s behaviour. Reinforcement learning is formulated as a gradient-based optimisation procedure, but reinforcement learning problems with unique optimal solutions are rare. This work focuses on ways to distinguish policies that solve the same reinforcement learning problem differently.
A pointed analysis of reinforcement learning literature, specifically of policy gradient methods, shows that policies are best understood through their behaviour: The actions that the policy selects based on the encountered state. There are two equivalent main approaches to understanding this behaviour empirically: Episodic behavioural characterisation, which interprets a policy as a random element in the space of state-action time series and behavioural characterisation through the occupancy measure, which interprets a policy as a random element representing the visitation frequency in the space of state-action pairs.
Both characterisations can be used to deduce general, problem-agnostic distance measures either by comparing time series or by comparing sampled distributions. Their utility is first explored by visualising small populations of policies. The relationships between policies that emerge in the visualisation also reflect other, hidden relationships, thereby validating the approach.
This idea is then integrated with an evolutionary algorithm. Evolutionary computation is a branch of optimisation techniques that improve populations of candidate solutions. In quality-diversity methods specifically, the members of the population are encouraged to differ from one another to improve the evolutionary process. When these methods are used to solve reinforcement learning problems, a measure of distance between candidate solutions is required. The developed techniques to distinguish policies are successfully connected with MAP-Elites, a quality-diversity heuristic that is often used for reinforcement learning problems and typically distinguishes policies based on selective, problem-specific criteria.
The applications validate the established general distances measure, but they also show that a diverse population of solutions is, for example, more explorative and more robust against environmental changes even if the diversity is not directed towards the specific goal such as exploration or robustness. This reveals the intrinsic value of diversity.},
url = {https://hdl.handle.net/20.500.11811/14349}
}
urn: https://nbn-resolving.org/urn:nbn:de:hbz:5-91296,
author = {{Arno Berengar Feiden}},
title = {Diversity in Reinforcement Learning},
school = {Rheinische Friedrich-Wilhelms-Universität Bonn},
year = 2026,
month = aug,
note = {Reinforcement learning is a machine learning paradigm that aims to develop optimal policies for discrete-time control problems based on quantitative feedback on the learner’s behaviour. Reinforcement learning is formulated as a gradient-based optimisation procedure, but reinforcement learning problems with unique optimal solutions are rare. This work focuses on ways to distinguish policies that solve the same reinforcement learning problem differently.
A pointed analysis of reinforcement learning literature, specifically of policy gradient methods, shows that policies are best understood through their behaviour: The actions that the policy selects based on the encountered state. There are two equivalent main approaches to understanding this behaviour empirically: Episodic behavioural characterisation, which interprets a policy as a random element in the space of state-action time series and behavioural characterisation through the occupancy measure, which interprets a policy as a random element representing the visitation frequency in the space of state-action pairs.
Both characterisations can be used to deduce general, problem-agnostic distance measures either by comparing time series or by comparing sampled distributions. Their utility is first explored by visualising small populations of policies. The relationships between policies that emerge in the visualisation also reflect other, hidden relationships, thereby validating the approach.
This idea is then integrated with an evolutionary algorithm. Evolutionary computation is a branch of optimisation techniques that improve populations of candidate solutions. In quality-diversity methods specifically, the members of the population are encouraged to differ from one another to improve the evolutionary process. When these methods are used to solve reinforcement learning problems, a measure of distance between candidate solutions is required. The developed techniques to distinguish policies are successfully connected with MAP-Elites, a quality-diversity heuristic that is often used for reinforcement learning problems and typically distinguishes policies based on selective, problem-specific criteria.
The applications validate the established general distances measure, but they also show that a diverse population of solutions is, for example, more explorative and more robust against environmental changes even if the diversity is not directed towards the specific goal such as exploration or robustness. This reveals the intrinsic value of diversity.},
url = {https://hdl.handle.net/20.500.11811/14349}
}





