课题基金 / 基金详情

Distributions of patterns and statistics in Markovian sequences

Distributions of patterns and statistics in Markovian sequences
马尔可夫序列中的模式和统计分布
批准号:
0805577
负责人:
Donald Martin
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-07-01 至 2013-06-30

项目摘要

项目成果

Donald Martin的其他基金

相似基金

相关文献

中文摘要
翻译
在这个项目中,研究者通过辅助马尔可夫链研究了序列中模式和统计分布的计算。在该方法中,利用原始序列中的马尔可夫结构将辅助马尔可夫链与序列关联起来,这样一来,当且仅当辅助马尔可夫链位于与该事件对应的一类状态中时,原始序列中的感兴趣事件才会发生。一旦辅助链建立起来,事件的概率就可以通过跟踪链中的运动来计算,然后提取期望的概率。这项工作的目标有三个:(1)计算迄今为止尚未解决的复杂模式的分布;(二)运用概率工具进行统计检验和数据分析;(3)对概率图模型建模的标记和分割数据统计中的不确定性进行量化。这些目标是集成的,因为计算模式和统计分布的概率方法为目标(2)和(3)的统计应用程序提供了必要的数学工具,而这些应用程序反过来又推动了在日益复杂的情况下计算分布的需求。虽然满足前两个目标将为文献提供重要贡献,但研究的主要贡献由目标(3)表示。隐状态序列统计抽样分布的计算提供了一种量化标记和分割数据不确定性的方法,这是一个尚未得到充分解决的领域。在对标记数据的统计推断感兴趣的情况下,典型的方法是确定给定观测值的最可能的状态序列,然后从该状态序列中获得感兴趣的统计值。然而,如果一个人对最好的标签集感兴趣,那么最有可能的状态是最优的,而对于标签的统计推断则可能不是如此。这项工作提供了一种新的方法来计算标记数据统计的精确抽样分布,为更准确的推理提供了一种手段。计算分布对估计参数的敏感性和对变化点的应用也将被考虑。对序列中与模式和统计相关的分布特性的需求,无论是从模型中产生的数据的实现,还是用于标记和分割观察数据的隐藏序列,都出现在许多具有大量数据集的实际研究领域,如生物信息学、时间序列、信息论、经济学、数据挖掘和质量控制。本研究开发了计算此类分布的计算工具。模式和统计分布的结果可应用于许多实际问题,例如检测DNA序列中的基因、启动子或其他具有重要功能的模式,以及确定与健康相关研究中观察分类有关的概率、表明经济数据新制度的变化点、表明入侵的模式或与监测工作有关的模式。该理论可用于计算被噪声或缺失观测值破坏的底层序列中的模式分布,也可用于计算通过组合或其他方法难以处理的统计分布。因此,这项研究促进了新的科学研究,这些研究依赖于迄今为止尚未计算过的模式或统计结果。
英文摘要
In this project the investigator studies the computation of distributions of patterns and statistics in sequences through auxiliary Markov chains. In the method, Markovian structure in the original sequence is exploited to associate an auxiliary Markov chain with the sequence in such a manner that an event of interest in the original sequence occurs if and only if the auxiliary Markov chain lies in a class of states that corresponds to the event. Once the auxiliary chain is set up, probabilities for the event may be computed by tracking movements through the chain and then extracting the desired probabilities. The goals of this work are threefold: (1) to compute distributions of complex patterns that have not been addressed to date; (2) to apply probabilistic tools that are developed to statistical testing and data analysis; and (3) to quantify uncertainty in statistics of labeled and segmented data modeled by probabilistic graphical models. These goals are integrated, since probabilistic approaches to computing distributions of patterns and statistics provide the mathematical tools necessary for the statistical applications of goals (2) and (3), and in turn those applications drive the need for computing distributions in increasingly complex situations. Whereas satisfying the first two goals will provide an important contribution to the literature, the major contribution of the research is represented by goal (3). The computation of sampling distributions of statistics of hidden state sequences provides a method of quantifying uncertainty in labeled and segmented data, an area that has not been adequately addressed. In cases where one is interested in inference on statistics of labeled data, a typical approach is to determine the most likely sequence of states given the observations, and then obtain the value of the statistic of interest from that state sequence. However, whereas the most likely states are optimal if one is interested in the best set of labels, it may not be so for inference on statistics of the labels. This work provides a novel approach to compute the exact sampling distribution of statistics of labeled data, providing a means for more accurate inference. Sensitivity of computed distributions to estimated parameters and applications to change points will also be considered. The need for distributional properties associated with patterns and statistics in sequences, both realizations of data emanating from a model and hidden sequences used to label and segment observed data, arises in many practical fields of study with massive data sets, such as bioinformatics, time series, information theory, economics, data mining, and quality control. In this research computational tools are developed for computing such distributions. Results for distributions of patterns and statistics may be applied to many practical problems, such as detecting genes, promoters, or other functionally significant patterns in DNA sequences, and determining probabilities related to classification of observations in health-related studies, change points that indicate new regimes in economic data, patterns that indicate an intrusion, or of patterns associated with surveillance work. The theory may be used to compute distributions of patterns in underlying sequences that are corrupted by noise or missing observations, and also distributions of statistics that are intractable by combinatorial or other means. Thus this research facilitates new scientific studies that rely on results for patterns or statistics that have not been computed to date.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical Analysis of Categorical Time Series through Sparse Markov Models
  • 批准号:
    1811933
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.0万
  • 财政年份:
    2018
  • 负责人:
    Donald Martin
  • 依托单位:
Distribution of Patterns and Statistics in Random Sequences
  • 批准号:
    1107084
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.0万
  • 财政年份:
    2011
  • 负责人:
    Donald Martin
  • 依托单位:
Urban Systemic Program in Science, Mathematics, and Technology Education (USP): SciMaX
  • 批准号:
    0114949
  • 项目类别:
    Cooperative Agreement
  • 资助金额:
    $0.0万
  • 财政年份:
    2001
  • 负责人:
    Donald Martin
  • 依托单位:
CPMSA: "Comprehensive Partnerships for Minority Student Achievement"
  • 批准号:
    9550622
  • 项目类别:
    Cooperative Agreement
  • 资助金额:
    $223.82万
  • 财政年份:
    1995
  • 负责人:
    Donald Martin
  • 依托单位:
海外基金