False-positive selection identified by ML-based methods:: Examples from the Sig1 gene of the diatom Thalassiosira weissflogii and the tax gene of a human T-cell lymphotropic virus

False-positive selection identified by ML-based methods:: Examples from the Sig1 gene of the diatom Thalassiosira weissflogii and the tax gene of a human T-cell lymphotropic virus
复制标题

DOI:
10.1093/molbev/msh098
复制
发表时间:
2004-05-01
影响因子:
10.7
通讯作者:
Nei, M
Nei, M
中科院分区:
生物学1区
文献类型:
--
作者:
Suzuki, Y;Nei, M

文献摘要

被引文献

相似文献

性诱导基因1(Sexually induced gene 1,Sig 1)是一种配子识别蛋白。Sorhannus(2003)使用简约分析和基于最大似然(ML)的贝叶斯方法分析了Sig 1的核苷酸序列,以推断单个氨基酸位点处的阳性选择,并报道了通过后一种方法检测到阳性选择位点,但通过前一种方法检测不到。然后他得出结论,对于这种类型的研究,基于ML的方法比简约分析更可靠。在这里,我们表明,他的结果显然代表假阳性的情况下,ML为基础的方法,并没有确凿的证据表明,该基因包含积极选择的网站。我们进一步证明,在人T细胞嗜淋巴细胞病毒I型(HTLV-I)的税收基因,所有的密码子的网站,包括不变的网站,可以推断为积极的选择网站的ML为基础的方法。这些观察结果表明,基于ML的方法可能会产生许多假阳性位点。假阳性发生的主要原因之一是在基于ML的方法中,密码子位点被分组为几个类别,在纯统计的基础上具有不同的非同义/同义比率(ω),并且通过检查每个类别的平均ω是否大于1来间接推断阳性选择。然而,在简约分析中,检查每个密码子位点处的核苷酸的进化变化。出于这个原因,基于简约的方法很少产生假阳性,并且比基于ML的方法更安全,用于检测单个密码子位点处的阳性选择,尽管需要大量的序列。
Sexually induced gene 1 (Sig1) in the centric diatom Thalassiosira weissflogii is considered to encode a gamete recognition protein. Sorhannus (2003) analyzed nucleotide sequences of Sig1 using parsimony analysis and the maximum-likelihood (ML)-based Bayesian method for inferring positive selection at single amino acid sites and reported that positively selected sites were detected by the latter method but not by the former. He then concluded that for this type of study, the ML-based method is more reliable than parsimony analysis. Here we show that his results apparently represent false-positive cases of the ML-based method and that there is no solid evidence that this gene contains positively selected sites. We further demonstrate that in the tax gene of human T-cell lymphotropic virus type I (HTLV-I), all codon sites, including invariable sites, can be inferred as positively selected sites by the ML-based method. These observations indicate that the ML-based method may produce many false-positive sites. One of the main reasons for the occurrence of false positives is that in the ML-based method, codon sites are grouped into several categories, with different nonsynonymous/synonymous rate ratios (omegas), on a purely statistical basis, and positive selection is inferred indirectly by examining whether the average omega for each category is greater than 1. In parsimony analysis, however, the evolutionary change of nucleotides at each codon site is examined. For this reason, parsimony-based methods rarely produce false positives and are safer than ML-based methods for detecting positive selection at individual codon sites, although a large number of sequences are necessary.