HPRS: hierarchical potential-based reward shaping from task specifications
{{output}}
The automatic synthesis of policies for robotics systems through reinforcement learning relies upon, and is intimately guided by, a reward signal. Consequently, this signal should faithfully reflect the designer's intentions, which are often expressed as a co... ...