Influence-aware memory architectures for deep reinforcement learning in POMDPs
{{output}}
Due to its perceptual limitations, an agent may have too little information about the environment to act optimally. In such cases, it is important to keep track of the action-observation history to uncover hidden state information. Recent deep reinforcement le... ...