A Reinforcement Learning Memristive Circuit Based on Q-Learning and Operant Conditioning
Junwei Sun, Lingying Kong, Yan He, Yanfeng Wang
Zhengzhou University of Light Industry
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
In most memristive neural network circuits based on operant conditioning, the agent’s tendency towards certain behaviors is simply reflected through changes in synaptic weight. No specific analysis has been conducted on the changes in the tendency of intelligent agent behavior. Therefore, an operant conditioning circuit based on memristors and Q-learning is proposed to analyze the behavior and decision-making of intelligent agents in complex environments. The designed network uses Q-learning to update the decision voltage in the circuit based on the optimal Bellman equation, allowing agents to make different strategies according to the constantly changing environment to achieve optimal results. In addition, this study also uses the Double Q-learning to effectively reduce the overestimation bias of Q-learning and improve the stability and strategy performance of agent learning by introducing Q value estimation. This design provides more References for the development of automobile obstacle avoidance systems.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
工程Advanced Memory and Neural Computing
Machine Learning and ELM · Neural Networks and Applications
参考文献 52
此处列出前 3 条
引用本文 1
按被引量排序,此处列出前 3 条