Alignment uncertainty and genomic analysis

Alignment uncertainty and genomic analysis
复制标题

DOI:
10.1126/science.1151532
复制
发表时间:
2008-01-25
期刊:
影响因子:
56.9
通讯作者:
Huelsenbeck, John P.
Huelsenbeck, John P.
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Wong, Karen M.;Suchard, Marc A.;Huelsenbeck, John P.

文献摘要

被引文献

相似文献

应用于基因组数据分析的统计方法没有考虑到序列比对的不确定性。实际上,这种对齐被视为一种观察,所有后续的推论都依赖于对齐是否正确。对于许多系统发育研究来说,这可能不是太大的问题,在这些研究中,基因是精心选择的,除其他外,容易对齐。然而,在比较基因组学研究中,同样的统计方法被反复应用于数千个基因,其中许多基因很难对齐。利用7种酵母菌的基因组数据,我们发现不确定的比对可能导致几个问题,包括不同的比对方法导致不同的结论。
The statistical methods applied to the analysis of genomic data do not account for uncertainty in the sequence alignment. Indeed, the alignment is treated as an observation, and all of the subsequent inferences depend on the alignment being correct. This may not have been too problematic for many phylogenetic studies, in which the gene is carefully chosen for, among other things, ease of alignment. However, in a comparative genomics study, the same statistical methods are applied repeatedly on thousands of genes, many of which will be difficult to align. Using genomic data from seven yeast species, we show that uncertainty in the alignment can lead to several problems, including different alignment methods resulting in different conclusions.