Computationally efficient estimation of the error rates of hidden Markov model results
Computationally efficient estimation of the error rates of hidden Markov model results
批准号:
0914739
负责人:
Patrick Van Roey
金额:
$30.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-08-01 至 2013-07-31
中文摘要
NSF提案:0914739PI:Newberg,Lee A.隐马尔可夫模型结果的错误率的计算高效估计隐马尔可夫模型被广泛应用于各种领域,包括语音识别、计量经济学、计算机视觉、信号处理、密码分析和计算生物学。在语音识别中,隐马尔可夫模型可用于根据声音的某些质量的时间序列将一个词与另一个词区分开来。在金融学中,这些模型可以用来模拟低、中、高债务违约制度之间的未知过渡。在计算机视觉中,它们可以用来解码美国手语(ASL)。隐马尔可夫模型在计算生物学中被用来发现核苷酸(DNA或RNA)或多肽(蛋白质)序列之间的相似性,并预测蛋白质的结构。隐马尔可夫模型的使用是因为它们允许简单地描述和实现强大的统计模型和算法,用于在序列数据中评分匹配可能性。也许隐马尔可夫模型最常见的用途是用于假设检验或分类。例如,语音识别模型可以用来量化认为录制的消息包含单词?大象的信念。然而,一旦一个信念的分数被计算出来,问题是如何解释这个值。分数是否强到足以指示信号,或者噪声产生如此强的分数的可能性是否合理?2.分数是否弱到足以指示噪声,或者信号产生如此弱的分数的可能性是否合理?分数阈值的假阳性率(与I类错误或p值密切相关)是噪声数据产生的分数至少与阈值一样强的概率。分数阈值的假阴性率是信号数据不能得到至少与阈值一样强的分数的概率。2008年,Newberg设计了一种估计错误率的方法,该方法比适用于一般隐马尔可夫模型的其他方法更有效。然而,这种方法对于计算密集型应用程序来说仍然太慢,例如重复搜索大型DNA数据库。这项研究旨在通过两种方法加快估计速度:(1)创造性地重复使用模拟,(2)统计稳健地消除模拟结果的离群值。这项研究具有重要意义,因为错误率的易得性允许各种科学领域的研究人员评估他们的结论的统计意义和他们的假设检验的力量。一旦技术和软件可用,语音识别或ASL识别的研究人员将能够使用严格推导的统计显著性值来设置他们的单词识别假设检验阈值。金融建模师将有一个严格的标准来评估他们的市场时机选择模型。计算生物学家将对他们的序列比对和他们对大型序列数据库的模式扫描具有严格的统计意义值。更一般地,隐马尔可夫模型结果的错误率的可用性将显著增强在当前未使用隐马尔可夫模型的领域中使用的隐马尔可夫模型的吸引力。
英文摘要
NSF proposal: 0914739PI: Newberg, Lee A.Computationally efficient estimation of the error rates of hidden Markov model resultsHidden Markov models are employed in a wide variety of fields, including speech recognition, econometrics, computer vision, signal processing, cryptanalysis, and computational biology. In speech recognition, hidden Markov models can be used to distinguish one word from another based upon the time series of certain qualities of a sound. In finance, the models can be used to simulate the unknown transitions between low, medium, and high debt default regimes in time. In computer vision they can be used to decode American Sign Language (ASL). Hidden Markov models are used in computational biology to find similarity between sequences of nucleotides (DNA or RNA) or polypeptides (proteins) and to predict protein structure.Hidden Markov models are employed because they permit the facile description and implementation of powerful statistical models and algorithms that are used to score a match possibility in sequence data. Perhaps the most common use of hidden Markov models is for the purpose of hypothesis testing or classification. For instance, a speech-recognition model may be used to quantify the belief that a recorded message contains the word ?elephant.? However, once a score for a belief has been computed, the question is how to interpret that value.1. Is the score strong enough to indicate a signal, or is it reasonably probable that noise will yield a score this strong?2. Is the score weak enough to indicate noise, or is it reasonably probable that a signal will yield a score this weak?The false positive rate (closely related to the type I error or p-value) for a score threshold is the probability that noise data will yield a score at least as strong as the threshold. The false negative rate for a score threshold is the probability that signal data will fail to score at least as strong as the threshold.In 2008, Newberg designed a method for estimating error rates that is more efficient than other approaches that are applicable to general hidden Markov models. However the approach is still too slow for computationally intensive applications such as repeated searches of large DNA databases. This proposed research aims to speed the estimation primarily via two approaches: (1) the creative re-use of simulations, and (2) statistically robust elimination of outlier simulation results.The proposed research is significant because the facile availability of error rates permits researchers in a wide variety of scientific fields to evaluate the statistical significance of their conclusions and the power of their hypothesis tests. Once the technique and software are available, researchers in speech recognition or ASL recognition will be able to use rigorously derived statistical significance values to set their hypothesis test thresholds for word recognition. Financial modelers will have a rigorous standard by which to evaluate their market timing models. Computational biologists will have rigorous statistical significance values for their sequence alignments and for their pattern scans of large sequence databases. More generally, the availability of error rates for hidden Markov model results will significantly enhance the attractiveness of hidden Markov models for use in fields where hidden Markov models are not currently employed.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
NANOSCALE: An Electronic Component from Protein Self-Assembly
-
批准号:9986431
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2000
-
负责人:Patrick Van Roey
-
依托单位:
国内基金
海外基金
固定参数可解算法在平面图问题的应用以及和整数线性规划的关系
-
批准号:60973026
-
项目类别:面上项目
-
资助金额:32.0万元
-
批准年份:2009
-
负责人:鲁道夫
-
依托单位: