Studies on optimal theory and its application in probabilistic decision processes with general utility functions.
Studies on optimal theory and its application in probabilistic decision processes with general utility functions.
批准号:
11640118
负责人:
KADOTA Yoshinobu
金额:
$0.9万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
1999
资助国家:
日本
项目状态:
已结题
起止时间:
1999 至 2000
中文摘要
停止决策过程是马尔可夫决策过程和停止问题的组合模型。MDP由可数状态集S、紧动作空间A(i)、转移概率q=(qa<ij>)和一致有界的立即报酬函数r(i,a,j)定义,它们对任意i,j ∈ S在a ∈ A(i)中连续.策略π是A(i_t)上的概率序列,由每个历史(i_0,a_0,i_1,.,i_t)(t= 0,1,.)约束.用σ表示停时,g表示效用函数,设B(t)= n ^t_<k=1> r(X_<k-1>,Δ_<k-1>,X_k),其中X_t和Δ_t分别是t时刻的状态和动作。(π,σ)对称为(i_0,α_0)-最优的,如果它使E^π_<i_0>[g(α_0+B(σ))]最大化,其中E^π_<i_0>[g(α_0+B(t))]是初始状态i_0在样本空间Ω=(S×A)^∞上的概率测度的期望。假设g是非减的、凹的且上有界的,或者g在真实的线R的任何紧子集上有界导数,对任意π,i,其中g^+是g的正部分。设v(i,α)= max_<{(π,σ)}> E^π_i(g(α+B(σ)).然后,我们得到了以下结果。1.对任意i ∈ S和α,α(i,α)满足最优性方程<$(i,α)= max {g(α),max_<α∈A><$_<j∈S> q_<ij>(a)<$(j,α+r(i,a,j)}(1)进一步,设(π,σ)满足P^π_<i_0>(σ>1)=1.2.若(π,σ)是(i_0,α_0)-最优对,则E^π_<i_0>[g(α_0 + B(σ))]满足(1).若E^π_<i_0>[g(α_0 + B(σ))]满足(1),则(π,σ)是(i_0,α_0)-最优的.
英文摘要
Stopped decision process is a combined model of Markov decision processes (MDPs) and the stopping problem. MDPs are specified by the set of countable states S, compact action space A (i) assigned at each state i ∈ S, transition probabilities q= (q_<ij>(a)), and a uniformly bounded immediate reward function r(i, a, j), which are continuous in a ∈ A(i) for any i, j ∈ S.A policy π is a sequence of probabilities on A (i_t) conditioned by each histories (i_0, a_0, i_1, …, i_t ) for t=0,1, …. Denote by σ a stopping time and by g a utility function.Let B(t)=Σ^t_<k=1> r (X_<k-1>, Δ_<k-1>, X_k), where X_t and Δ_t are the state and action at time t, respectively. The pair (π, σ) is called (i_0, α_0)-optimal if it maximizes E^π_<i_0> [g (α_0+B(σ))], where E^π_<i_0> is the expectation by the probability measure on the sample space Ω=(S×A)^∞ for an initial state i_0.It is assumed that g is non-decreasing, concave and bounded above, or that g has an bounded derivative on any compact subset of the real line R satisfying E^π_i [sup_<t【greater than or equal】0> g^+(α_0+B(t))] < ∞ for any π, i, where g^+ is the positive part of g. Let v(i, α) = max_<{(π, σ)}> E^π_i (g(α+B(σ)). Then, we have following results.1. For any i ∈ S and α, υ(i, α) satisfies optimality equationsυ(i, α) = max {g(α), max_<α∈A> Σ_<j∈S> q_<ij>(a) υ (j, α+r(i, a, j)}(1)Furthermore, suppose (π, σ) satisfies P^π_<i_0> (σ>1)=1.2. If (π, σ) is (i_0, α_0)-optimal pair, then E^π_<i_0> [g(α_0 + B(σ))] satisfies (1).3. If E^π_<i_0>[g(α_0 + B(σ))] satisfies (1), then (π, σ) is (i_0, α_0)-optimal.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Y.Kadota: "Deviation matrix, Laurent series and Blackwell optimality in countable state Markov decision Processes."To appear. Lecture note in Institute of Math. Anal. in Kyoto Univ.. (2001)
Y.Kadota:“可数状态马尔可夫决策过程中的偏差矩阵、洛朗级数和布莱克威尔最优性。”出现。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Y.Kadota, M.Kurano and M.Yasuda: "Risk-averse stopped Markov decision processes."The 4th BIC (Bull. Inform. & Cybernet.) symposium. (1999)
Y.Kadota、M.Kurano 和 M.Yasuda:“风险规避阻止了马尔可夫决策过程。”第四个 BIC(Bull.Inform.
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Y.Kadota, M.Kurano and M.Yasuda: "Stopped decision processes in conjunction with general utility."To appear in J.Inform. & Optim.Sci.. (2001)
Y.Kadota、M.Kurano 和 M.Yasuda:“停止与一般实用程序相关的决策过程。”出现在 J.Inform 中。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Y.Kadota.: "Deviation matrix,Laurent series and Blackwell optimality in countable state Markov decision processes."数理解析研究所講究録「不確実なモデルによる動的計画理論の課題とその展望」. (掲載予定)(某雑誌に掲載予定). (2001)
Y. Kadota.:“可数状态马尔可夫决策过程中的偏差矩阵、洛朗级数和布莱克威尔最优性。”数学科学研究所《不确定模型动态规划理论的问题与展望》(待出版)(预定)发表在某杂志上)(2001)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Y.Kadota,M.Kurano and M.Yasuda.: "Stopped decision processes in conjunction with general utility."To appear in J.Information and Optimization Science.. (2001)
Y.Kadota、M.Kurano 和 M.Yasuda.:“停止与通用实用程序相关的决策过程。”出现在 J.Information and Optimization Science.. (2001)
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 8 条
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
补偿性还是非补偿性规则:探析风险决策的行为与神经机制
-
批准号:31170976
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2011
-
负责人:李纾
-
依托单位:
基于神经营销学方法的品牌延伸认知与决策研究
-
批准号:70772048
-
项目类别:面上项目
-
资助金额:20.0万元
-
批准年份:2007
-
负责人:马庆国
-
依托单位: