首页 正文

HPRS: hierarchical potential-based reward shaping from task specifications

{{output}}
The automatic synthesis of policies for robotics systems through reinforcement learning relies upon, and is intimately guided by, a reward signal. Consequently, this signal should faithfully reflect the designer's intentions, which are often expressed as a co... ...