Markov decision processes
Eitan Altman
内容与影响
Markov decision processes (MDPs), also known as controlled Markov chains, constitute a basic framework for dynamically controlling systems that evolve in a stochastic way. MDPs are thus a generalization of (non-controlled) Markov chains, and many useful properties of Markov chains carry over to controlled Markov chains. In deterministic models, where the transition probabilities are only zero or one, the controller can fully predict the evolution of the state of the system as a result of applying a sequence of actions, if it knows the initial state. This chapter concludes that the Markov policies are sufficiently rich so that a cost that can be achieved by an arbitrary policy can also be achieved by a Markov policy.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
回答优先基于摘要、文献信息与可获取全文;依据不足时会明确说明。
学术脉络
学科主题
计算机 / AIReinforcement Learning in Robotics
Advanced Queuing Theory Analysis · Markov Chains and Monte Carlo Methods