HPRS: hierarchical potential-based reward shaping from task specifications

The automatic synthesis of policies for robotics systems through reinforcement learning relies upon, and is intimately guided by, a reward signal. Consequently, this signal should faithfully reflect the designer's intentions, which are often expressed as a co... ...

请注册登录后继续浏览