Distribution of Patterns and Statistics in Random Sequences
Distribution of Patterns and Statistics in Random Sequences
批准号:
1107084
负责人:
Donald Martin
金额:
$10.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-08-01 至 2015-09-30
中文摘要
该项目的目标是开发和应用工具,用于有效计算随机序列中与模式和统计相关的分布,既实现观测序列,也实现以观测数据为条件的隐藏状态序列。由于可用的数据集规模巨大,不仅准确而且高效的计算方法非常重要。拟议的研究旨在为这一领域作出贡献。最小确定性有限自动机、概率生成函数、和积算法和矩阵向量更新是用来形成有效计算模式和统计分布的算法的主要工具。本研究的目标有三个:(i)进一步开发有效的方法来量化概率图模型中隐藏状态序列的统计不确定性;(ii)有效地精确计算以前没有计算过的复杂模式的概率分布;(三)将(一)和(二)的概率工具应用于统计检验和数据分析。隐藏状态统计中的不确定性量化通常通过确定特定标准的最优状态序列来处理,然后以确定性的方式从最优状态序列中评估统计量。然而,这种方法并不能解释各州的不确定性。另一种方法是从给定观察数据的状态的条件分布中抽样,然后根据经验近似统计量的分布。然而,为了使近似准确,需要许多样本,从而导致可伸缩性问题。我们给出了一种有效计算精确分布的方法。项目的目标是集成的,因为应用程序中的统计推断需要模式和统计分布,而这些应用程序反过来又在日益复杂的情况下驱动对计算分布的需求。这项工作的目标包括计算蛋白质-蛋白质相互作用预测错误率的基于模型的分布,计算DNA序列同源性搜索的间隔种子覆盖率的精确分布,以及计算多状态高阶马尔可夫试验的一维扫描统计量的精确分布。对随机序列中与模式和统计相关的分布特性的需求出现在许多具有大量数据集的实际研究领域,如生物信息学、时间序列、信息论、经济学、数据挖掘和质量控制。本研究开发了计算此类分布的计算工具。模式和统计分布的结果可应用于许多实际问题,例如检测DNA序列中的基因、启动子或其他具有重要功能的模式,确定与健康相关研究中观察分类有关的概率,确定经济数据中表明新制度的变化点,确定表明计算机系统被入侵的模式,或与监视工作有关的模式。该理论将用于计算难以通过组合或其他方式处理的统计分布,并在通常通过模拟许多数据集来处理的情况下以有效的方式提供精确的概率。因此,这项研究促进了新的科学研究,这些研究依赖于与迄今尚未计算过的模式或统计数据相关的分布结果。
英文摘要
This project targets the development and application of tools for efficient computation of distributions associated with patterns and statistics in random sequences, both realizations of observation sequences as well as hidden state sequences conditional on observed data. Due to the massive size of data sets that are available, computational methods that are not only accurate but also efficient are important. The proposed research seeks to contribute in this area. Minimal deterministic finite automata, probability generating functions, the sum-product algorithm and matrix-vector updates are the primary tools used to form algorithms for efficiently computing distributions of patterns and statistics. The goals of this research are threefold: (i) to further develop efficient methods for quantifying uncertainty in statistics of hidden state sequences of probabilistic graphical models; (ii) to efficiently compute exact probability distributions of complex patterns that have not previously been computed; (iii) to apply the probabilistic tools of (i) and (ii) to statistical tests and data analysis. The quantification of uncertainty in statistics of hidden states is frequently dealt with by determining the state sequence that is optimal for a particular criterion, with the statistic then evaluated from the optimal state sequence in a deterministic fashion. However, that approach does not account for uncertainty in the states. An alternate approach is to sample from the conditional distribution of states given the observed data, and then approximate the distribution of the statistic empirically. However, many samples are needed so that the approximation is accurate, leading to problems with scalability. We give a way to compute exact distributions in an efficient manner. The goals of the project are integrated, in that distributions of patterns and statistics are needed for statistical inference in applications, and in turn those applications drive the need for computing distributions in increasingly complex situations. Objectives of the work include computing a model-based distribution of prediction error rates for protein-protein interactions, computing exact distributions of coverage of spaced seeds for homology searches in DNA sequences, and computing the exact distribution of the one-dimensional scan statistic for multi-state higher-order Markovian trials. The need for distributional properties associated with patterns and statistics in random sequences arises in many practical fields of study with massive data sets, such as bioinformatics, time series, information theory, economics, data mining, and quality control. In this research computational tools are developed for computing such distributions. Results for distributions of patterns and statistics may be applied to many practical problems, such as detecting genes, promoters, or other functionally significant patterns in DNA sequences, determining probabilities related to classifications of observations in health-related studies, change points that indicate new regimes in economic data, patterns that indicate an intrusion in a computer system, or patterns associated with surveillance work. The theory will be used to compute distributions of statistics that are intractable by combinatorial or other means, and to provide exact probabilities in an efficient manner in situations that are typically handled by simulation of many data sets. Thus this research facilitates new scientific studies that rely on results for distributions associated with patterns or statistics that have not been computed to date.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical Analysis of Categorical Time Series through Sparse Markov Models
-
批准号:1811933
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2018
-
负责人:Donald Martin
-
依托单位:
Distributions of patterns and statistics in Markovian sequences
-
批准号:0805577
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2008
-
负责人:Donald Martin
-
依托单位:
Urban Systemic Program in Science, Mathematics, and Technology Education (USP): SciMaX
-
批准号:0114949
-
项目类别:Cooperative Agreement
-
资助金额:$0.0万
-
财政年份:2001
-
负责人:Donald Martin
-
依托单位:
CPMSA: "Comprehensive Partnerships for Minority Student Achievement"
-
批准号:9550622
-
项目类别:Cooperative Agreement
-
资助金额:$223.82万
-
财政年份:1995
-
负责人:Donald Martin
-
依托单位:
Mathematical Sciences: Recursion Theory and Set Theory
-
批准号:9505153
-
项目类别:Continuing Grant
-
资助金额:$24.0万
-
财政年份:1995
-
负责人:Donald Martin
-
依托单位:
Mathematical Sciences: Recursion Theory and Set Theory
-
批准号:9206946
-
项目类别:Continuing Grant
-
资助金额:$39.7万
-
财政年份:1992
-
负责人:Donald Martin
-
依托单位:
Mathematical Sciences: Recursion Theory and Set Theory
-
批准号:8902555
-
项目类别:Continuing Grant
-
资助金额:$38.94万
-
财政年份:1989
-
负责人:Donald Martin
-
依托单位:
Mini-Computer Applications to Undergraduate Meteorology Instruction
-
批准号:7813134
-
项目类别:Standard Grant
-
资助金额:$1.53万
-
财政年份:1978
-
负责人:Donald Martin
-
依托单位:
海外基金