Enhancement of Hippocampal Spatial Decoding with a Q-Learning Method Utilizing Theta Rhythm Phase Precession as the Relative Reward Approach
收藏资源简介:
Winners of the 2014 Nobel Prize in Physiology or Medicine, Professors John O’Keefe, May‐Britt Moser and Edvard I. Moser found that the internal global positioning system (GPS) in the brain allows us to be able to flexibly navigate the world they live in – exploring new areas, returning quickly to remembered places, and taking shortcuts and confirmed that place cells in hippocampus and grid cells in entorhinal cortex (EC) are responsible for higher-order cognitive map of the environment. Indeed, these abilities feel so easy and natural that it is not immediately obvious how complex the underlying processes really are. In contrast, spatial navigation remains a substantial challenge for artificial agents whose abilities are far outstripped by those of mammals. Hippocampal place cells and interneurons in mammals have proved that they own stable place fields and theta phase precession profiles to encode the spatial information from the environment. The hippocampal CA1 neurons can be represented as the location of the animal and the prospective information of goal location. Reinforcement learning algorithm, e.g., Q-learning, has been adopted to build a navigation model of place cells for the purpose of addressing goal direction navigation problems.<br> In this study, we propose dynamical Q-learning (dQ-learning), because of its adaptive reward function based on theta phase precession, which has recently been associated with a rat’s experiences at destinations, and use of information from both place cells and interneurons as inputs to predict the animal’s trajectory. We evaluated the convergence rates and learning performances of tQ-learning and dQ-learning with different cell types. The results demonstrate that dQ-learning improves learning performance and convergence rate and place cells and interneurons with phase precession may provide valuable information to improve the prediction of trajectory. To investigate whether the enhancement of hippocampal spatial decoding with the dQ-learning method was effective in goal-direction navigation, experimental data were recorded from rats implanted with microelectrodes and trained in a water reward task. During the task electrophysiological recordings of spikes, LFPs, and movement trajectories were acquired. The proposed dQ-learning algorithm achieved better learning performance with good prediction accuracy and a high convergence rate. The adaptive reward function and cell types were found to be critical factors for hippocampal spatial decoding using the dQ-learning method.
2014年诺贝尔生理学或医学奖得主约翰·奥基夫(John O’Keefe)、梅-布里特·莫泽(May‐Britt Moser)与爱德华·I·莫泽(Edvard I. Moser)教授发现,大脑内部的全球定位系统(GPS)可让我们灵活穿梭于所处的世界——探索陌生区域、快速返回记忆中的地点、抄近路;同时他们证实,海马体内的位置细胞与内嗅皮层(entorhinal cortex, EC)内的网格细胞,负责构建环境的高阶认知地图。事实上,这类行为看似轻松自然,人们往往难以立刻洞悉其背后的复杂机制。与之形成鲜明反差的是,空间导航对于人工智能体(AI Agent)而言仍是一项重大难题,其能力远不及哺乳动物。哺乳动物的海马体位置细胞与中间神经元已被证实具备稳定的位置野与θ相位进动(theta phase precession)特征,可用于编码环境中的空间信息。海马CA1神经元能够表征动物的当前位置与目标位置的预期信息。强化学习(Reinforcement Learning)算法(如Q学习(Q-learning))已被用于构建基于位置细胞的导航模型,以解决目标导向导航问题。 本研究提出动态Q学习(dynamical Q-learning, dQ-learning)算法,其基于θ相位进动设计自适应奖励函数——该机制近期被发现与大鼠在目的地的经历相关——并同时利用位置细胞与中间神经元的信息作为输入,以预测动物的运动轨迹。我们针对不同细胞类型,对比了传统Q学习(tQ-learning)与动态Q学习的收敛速度与学习性能。实验结果表明,动态Q学习可提升学习性能与收敛速度,且具备相位进动特性的位置细胞与中间神经元,可为优化轨迹预测提供有价值的信息。为验证基于动态Q学习的海马体空间解码增强方法是否可有效应用于目标导向导航任务,研究团队记录了植入微电极的大鼠在水奖励任务中的实验数据。任务期间,同步采集了锋电位(spikes)、局部场电位(Local Field Potentials, LFPs)与运动轨迹的电生理记录数据。所提出的动态Q学习算法实现了更优异的学习性能,兼具良好的预测精度与较高的收敛速度。研究发现,自适应奖励函数与细胞类型是基于动态Q学习方法开展海马体空间解码的关键影响因素。



