A model of the statistical power of comparative genome sequence analysis

A model of the statistical power of comparative genome sequence analysis
复制标题

DOI:
10.1371/journal.pbio.0030010
复制
发表时间:
2005-01-01
期刊:
影响因子:
9.8
通讯作者:
Eddy, SR
Eddy, SR
中科院分区:
生物学1区
文献类型:
--
作者:
Eddy, SR

文献摘要

被引文献

相似文献

比较基因组序列分析是强大的,但测序基因组是昂贵的。希望能够预测比较基因组学需要多少基因组,以及进化距离是多少。在这里,我描述了一个简单的数学模型的共同问题,确定保守序列。该模型导致一些有用的经验法则。对于一个给定的进化距离,在确定保守区域的统计严格性的恒定水平所需的比较基因组的数量与待检测的保守特征的大小成反比。在短的进化距离,所需的比较基因组的数量也与距离成反比。这些缩放行为为未来的比较基因组测序需求提供了一些直觉,例如使用密切相关的比较基因组的“系统发育阴影”方法的建议使用,以及小保守特征的高分辨率检测的可行性。
Comparative genome sequence analysis is powerful, but sequencing genomes is expensive. It is desirable to be able to predict how many genomes are needed for comparative genomics, and at what evolutionary distances. Here I describe a simple mathematical model for the common problem of identifying conserved sequences. The model leads to some useful rules of thumb. For a given evolutionary distance, the number of comparative genomes needed for a constant level of statistical stringency in identifying conserved regions scales inversely with the size of the conserved feature to be detected. At short evolutionary distances, the number of comparative genomes required also scales inversely with distance. These scaling behaviors provide some intuition for future comparative genome sequencing needs, such as the proposed use of " phylogenetic shadowing'' methods using closely related comparative genomes, and the feasibility of high- resolution detection of small conserved features.