单调多边形
单调函数
马尔可夫决策过程
数学优化
数学
有界函数
马尔可夫过程
状态空间
贝尔曼方程
马尔可夫链
部分可观测马尔可夫决策过程
应用数学
决策问题
功能(生物学)
简单(哲学)
数理经济学
统计
算法
几何学
进化生物学
生物
认识论
数学分析
哲学
出处
期刊:Operations Research
[Institute for Operations Research and the Management Sciences]
日期:1987-10-01
卷期号:35 (5): 736-743
被引量:143
标识
DOI:10.1287/opre.35.5.736
摘要
This paper provides sufficient conditions for the optimal value in a discrete-time, finite, partially observed Markov decision process to be monotone on the space of state probability vectors ordered by likelihood ratios. The paper also presents sufficient conditions for the optimal policy to be monotone in a simple machine replacement problem, and, in the general case, for the optimal policy to be bounded from below by an easily calculated monotone function.
科研通智能强力驱动
Strongly Powered by AbleSci AI