Influence-aware memory architectures for deep reinforcement learning in POMDPs [0.03%]
基于影响感知的内存架构在POMDP中的深度强化学习研究
Miguel Suau,Jinke He,Elena Congeduti et al.
Miguel Suau et al.
Due to its perceptual limitations, an agent may have too little information about the environment to act optimally. In such cases, it is important to keep track of the action-observation history to uncover hidden state information. Recent d...
Jacopo Castellini,Sam Devlin,Frans A Oliehoek et al.
Jacopo Castellini et al.
Policy gradient methods have become one of the most popular classes of algorithms for multi-agent reinforcement learning. A key challenge, however, that is not addressed by many of these methods is multi-agent credit assignment: assessing a...
A maintenance planning framework using online and offline deep reinforcement learning [0.03%]
一种结合在线和离线深度强化学习的维护计划框架
Zaharah A Bukhsh,Hajo Molegraaf,Nils Jansen
Zaharah A Bukhsh
Cost-effective asset management is an area of interest across several industries. Specifically, this paper develops a deep reinforcement learning (DRL) solution to automatically determine an optimal rehabilitation policy for continuously de...
Ahmad Zainul Ihsan,Said Fathalla,Stefan Sandfeld
Ahmad Zainul Ihsan
The research in Materials Science and Engineering focuses on the design, synthesis, properties, and performance of materials. An important class of materials that is widely investigated are crystalline materials, including metals and semico...
Neuroevolution gives rise to more focused information transfer compared to backpropagation in recurrent neural networks [0.03%]
神经进化产生的递归神经网络信息传输相比反向传播更为集中
Arend Hintze,Christoph Adami
Arend Hintze
Artificial neural networks (ANNs) are one of the most promising tools in the quest to develop general artificial intelligence. Their design was inspired by how neurons in natural brains connect and process, the only other substrate to harbo...
Multi-objective reward generalization: improving performance of Deep Reinforcement Learning for applications in single-asset trading [0.03%]
多目标奖励泛化:改进深度强化学习在单一资产交易应用中的性能
Federico Cornalba,Constantin Disselkamp,Davide Scassola et al.
Federico Cornalba et al.
We investigate the potential of Multi-Objective, Deep Reinforcement Learning for stock and cryptocurrency single-asset trading: in particular, we consider a Multi-Objective algorithm which generalizes the reward functions and discount facto...
An Abstract Parabolic System-Based Physics-Informed Long Short-Term Memory Network for Estimating Breath Alcohol Concentration from Transdermal Alcohol Biosensor Data [0.03%]
基于抽象抛物线系统的物理信息长短期记忆网络:从经皮酒精生物传感器数据估计呼吸酒精浓度
Clemens Oszkinat,Susan E Luczak,I Gary Rosen
Clemens Oszkinat
The problem of estimating breath alcohol concentration based on transdermal alcohol biosensor data is considered. Transdermal alcohol concentration provides a promising alternative to classical methods such as breathalyzers or drinking diar...
Region-based evidential deep learning to quantify uncertainty and improve robustness of brain tumor segmentation [0.03%]
基于区域的证据深度学习量化不确定性并提高脑肿瘤分割的鲁棒性
Hao Li,Yang Nan,Javier Del Ser et al.
Hao Li et al.
Despite recent advances in the accuracy of brain tumor segmentation, the results still suffer from low reliability and robustness. Uncertainty estimation is an efficient solution to this problem, as it provides a measure of confidence in th...
Communicative capital: a key resource for human-machine shared agency and collaborative capacity [0.03%]
沟通资本:人机共因性和协作能力的关键资源
Kory W Mathewson,Adam S R Parker,Craig Sherstan et al.
Kory W Mathewson et al.
In this work, we present a perspective on the role machine intelligence can play in supporting human abilities. In particular, we consider research in rehabilitation technologies such as prosthetic devices, as this domain requires tight cou...
Knowledge- and ambiguity-aware robot learning from corrective and evaluative feedback [0.03%]
知识和模糊性感知的机器人学习:基于纠正性和评价性的反馈
Carlos Celemin,Jens Kober
Carlos Celemin
In order to deploy robots that could be adapted by non-expert users, interactive imitation learning (IIL) methods must be flexible regarding the interaction preferences of the teacher and avoid assumptions of perfect teachers (oracles), whi...