研究论文
Risk-Sensitive Markov Decision Processes
Ronald A. Howard, James E. Matheson
Stanford University SRI International
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
摘要 · 完整
This paper considers the maximization of certain equivalent reward generated by a Markov decision process with constant risk sensitivity. First, value iteration is used to optimize possibly time-varying processes of finite duration. Then a policy iteration procedure is developed to find the stationary policy with highest certain equivalent gain for the infinite duration case. A simple example demonstrates both procedures.
逐年被引趋势
47240
17
18
19
20
4721
22
23
24
25
26
关键指标
520
被引次数 · OpenAlex
2.69
领域内被引倍数
同类平均 = 1
同类平均 = 1
前 10%
引用位次
同领域 · 同年份 · 同类型
同领域 · 同年份 · 同类型
1
参考文献
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
论文问答
当前基于摘要回答
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AISimulation Techniques and Applications
Supply Chain and Inventory Management
参考文献 1
Dynamic Programming and Markov Processes.
被引 3,443Marshall L. Freimer, Ronald A. Howard · Journal of the American Statistical Association · 1961
此处列出前 3 条
引用本文 520
A comprehensive survey on safe reinforcement learning
被引 1,191Javier García, Fernando Fernández · Journal of Machine Learning Research · 2015
Risk-sensitive linear/quadratic/gaussian control
被引 499Peter J. L. Whittle · Advances in Applied Probability · 1981
On Foraging Time Allocation in a Stochastic Environment
被引 457Thomas B. Caraco · Ecology · 1980
按被引量排序,此处列出前 3 条