Combining evidence using p-values: application to sequence homology searches

Combining evidence using p-values: application to sequence homology searches
复制标题

DOI:
10.1093/bioinformatics/14.1.48
复制
发表时间:
1998-01-01
期刊:
影响因子:
5.8
通讯作者:
Gribskov, M
Gribskov, M
中科院分区:
生物学3区
文献类型:
--
作者:
Bailey, TL;Gribskov, M

文献摘要

被引文献

相似文献

动机:阐述一种直观且在统计学上有效的方法,用于合并独立的证据来源,从而为完整的证据得出一个p值,并将其应用于在序列同源性搜索中检测对多个模式的同时匹配问题。 结果:在序列分析中,对于一个序列(或序列区域)属于某一类别的成员关系,常常可以得到两个或更多(近似)独立的度量。鉴于所有可用的证据,我们希望估计该序列属于该类别的可能性。一个例子是估计一个大分子序列(DNA或蛋白质)与一组表征一个生物序列家族的模式(基序)的观察到的匹配的显著性。一种直观的做法是将每一项证据表示为一个p值,然后使用这些p值的乘积作为在该家族中的成员关系的度量。我们推导出一个公式和算法(QFAST),用于计算n个独立p值乘积的统计分布。我们证明,通过这个p值对序列进行排序有效地结合了多个基序中存在的信息,从而实现高度准确和灵敏的序列同源性搜索。
Motivation: To illustrate an intuitive and statistically valid method for combining independent sources of evidence that yields a p-value for the complete evidence, and to apply it to the problem of detecting simultaneous matches to multiple patterns in sequence homology searches.Results: In sequence analysis, two or more (approximately) independent measure of the membership of a sequence (or sequence region) in some class are often available. We would like to estimate the likelihood of the sequence being a member of the class in view of all the available evidence. an example is estimating the significance of the observed match of a macromolecular sequence (DNA or protein) to a set of patterns (motifs) that characterize a biological sequence family. An intuitive way to do this is to express each piece of evidence a as p-value, and then use the product of these p-values as the measure of membership in the family. We derive a formula and algorithm (QFAST) for calculating the statistical distribution of the product of n independent p-values. We demonstrate that sorting sequences by this p-value effectively combines the information present in multiple motifs, leading to highly accurate and sensitive sequence homology searches.