Studies on optimal theory and its application in probabilistic decision processes with general utility functions.
Studies on optimal theory and its application in probabilistic decision processes with general utility functions.
批准号:
11640118
负责人:
KADOTA Yoshinobu
金额:
$0.9万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
1999
资助国家:
日本
项目状态:
已结题
起止时间:
1999 至 2000
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Stopped decision process is a combined model of Markov decision processes (MDPs) and the stopping problem. MDPs are specified by the set of countable states S, compact action space A (i) assigned at each state i ∈ S, transition probabilities q= (q_<ij>(a)), and a uniformly bounded immediate reward function r(i, a, j), which are continuous in a ∈ A(i) for any i, j ∈ S.A policy π is a sequence of probabilities on A (i_t) conditioned by each histories (i_0, a_0, i_1, …, i_t ) for t=0,1, …. Denote by σ a stopping time and by g a utility function.Let B(t)=Σ^t_<k=1> r (X_<k-1>, Δ_<k-1>, X_k), where X_t and Δ_t are the state and action at time t, respectively. The pair (π, σ) is called (i_0, α_0)-optimal if it maximizes E^π_<i_0> [g (α_0+B(σ))], where E^π_<i_0> is the expectation by the probability measure on the sample space Ω=(S×A)^∞ for an initial state i_0.It is assumed that g is non-decreasing, concave and bounded above, or that g has an bounded derivative on any compact subset of the real line R satisfying E^π_i [sup_<t【greater than or equal】0> g^+(α_0+B(t))] < ∞ for any π, i, where g^+ is the positive part of g. Let v(i, α) = max_<{(π, σ)}> E^π_i (g(α+B(σ)). Then, we have following results.1. For any i ∈ S and α, υ(i, α) satisfies optimality equationsυ(i, α) = max {g(α), max_<α∈A> Σ_<j∈S> q_<ij>(a) υ (j, α+r(i, a, j)}(1)Furthermore, suppose (π, σ) satisfies P^π_<i_0> (σ>1)=1.2. If (π, σ) is (i_0, α_0)-optimal pair, then E^π_<i_0> [g(α_0 + B(σ))] satisfies (1).3. If E^π_<i_0>[g(α_0 + B(σ))] satisfies (1), then (π, σ) is (i_0, α_0)-optimal.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Y.Kadota: "Deviation matrix, Laurent series and Blackwell optimality in countable state Markov decision Processes."To appear. Lecture note in Institute of Math. Anal. in Kyoto Univ.. (2001)
Y.Kadota:“可数状态马尔可夫决策过程中的偏差矩阵、洛朗级数和布莱克威尔最优性。”出现。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Y.Kadota, M.Kurano and M.Yasuda: "Risk-averse stopped Markov decision processes."The 4th BIC (Bull. Inform. & Cybernet.) symposium. (1999)
Y.Kadota、M.Kurano 和 M.Yasuda:“风险规避阻止了马尔可夫决策过程。”第四个 BIC(Bull.Inform.
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Y.Kadota, M.Kurano and M.Yasuda: "Stopped decision processes in conjunction with general utility."To appear in J.Inform. & Optim.Sci.. (2001)
Y.Kadota、M.Kurano 和 M.Yasuda:“停止与一般实用程序相关的决策过程。”出现在 J.Inform 中。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Y.Kadota.: "Deviation matrix,Laurent series and Blackwell optimality in countable state Markov decision processes."数理解析研究所講究録「不確実なモデルによる動的計画理論の課題とその展望」. (掲載予定)(某雑誌に掲載予定). (2001)
Y. Kadota.:“可数状态马尔可夫决策过程中的偏差矩阵、洛朗级数和布莱克威尔最优性。”数学科学研究所《不确定模型动态规划理论的问题与展望》(待出版)(预定)发表在某杂志上)(2001)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Y.Kadota,M.Kurano and M.Yasuda.: "Stopped decision processes in conjunction with general utility."To appear in J.Information and Optimization Science.. (2001)
Y.Kadota、M.Kurano 和 M.Yasuda.:“停止与通用实用程序相关的决策过程。”出现在 J.Information and Optimization Science.. (2001)
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 8 条
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
补偿性还是非补偿性规则:探析风险决策的行为与神经机制
-
批准号:31170976
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2011
-
负责人:李纾
-
依托单位:
基于神经营销学方法的品牌延伸认知与决策研究
-
批准号:70772048
-
项目类别:面上项目
-
资助金额:20.0万元
-
批准年份:2007
-
负责人:马庆国
-
依托单位: