强化学习
背景(考古学)
人工智能
过程(计算)
计算机科学
口译(哲学)
差异(会计)
工作记忆
心理学
机器学习
认知心理学
联想学习
结合属性
认知
内容寻址存储器
计算
习惯
序列学习
认知科学
情景记忆
上下文模型
计算模型
学习理论
标识
DOI:10.1038/s41562-025-02340-0
摘要
Abstract Reinforcement learning (RL) algorithms have had tremendous success accounting for reward-based learning across species, including instrumental learning in contextual bandit tasks, and they capture variance in brain signals. However, reward-based learning in humans recruits multiple processes, including memory and choice perseveration; their contributions can easily be mistakenly attributed to RL computations. Here I investigate how much of reward-based learning behaviour is supported by RL computations in a context where other processes can be factored out. Reanalysis and computational modelling of 7 datasets ( n = 594) in diverse samples show that in this instrumental context, reward-based learning is best explained by a combination of a fast working-memory-based process and a slower habit-like associative process, neither of which can be interpreted as a standard RL-like algorithm on its own. My results raise important questions for the interpretation of RL algorithms as capturing a meaningful process across brain and behaviour.
科研通智能强力驱动
Strongly Powered by AbleSci AI