A probabilistic classifier for olfactory receptor pseudogenes

A probabilistic classifier for olfactory receptor pseudogenes
复制标题

DOI:
10.1186/1471-2105-7-393
复制
发表时间:
2006-08-29
期刊:
影响因子:
3
通讯作者:
Lancet, Doron
Lancet, Doron
中科院分区:
生物学4区
文献类型:
--
作者:
Menashe, Idan;Aloni, Ronny;Lancet, Doron

文献摘要

被引文献

相似文献

背景:嗅觉受体(Olfactory receptors,ORs)是哺乳动物中最大的基因超家族(900-1400个基因),在人类中含有50%以上的假基因。虽然大多数这些失活基因是通过编码框(无义)破坏来鉴定的,但看似完整的基因也可能由于其他有害的(错义)突变而失活。因此,最终评估的实际大小的功能的人类OR库需要一个准确的区分基因和pseudogenes.Results:为了表征不活跃的OR与完整的开放阅读框架,我们已经开发了一个概率分类嗅觉受体假基因(CORP)。该算法是基于偏离一个功能上至关重要的共识,构成60个高度保守的位置,通过比较两个进化约束的或剧目(小鼠和狗)与一个小的假基因部分。我们使用逻辑回归分析为保守位置分配适当的系数,从而实现活性和非活性OR之间的最大分离。因此,该算法仅将5%的小鼠功能性OR识别为假基因,将假阳性检测的上限设定为0.05。最后,我们使用该算法对384个据称完整的人类OR基因进行分类。其中,135被预测为可能编码非功能性蛋白质,和38个分离之间的活性和非活性形式由于错义polymorphisms.Conclusion:我们证明了CORP算法是能够区分功能和非功能性OR基因具有高精度,即使编码的蛋白质将由一个单一的氨基酸不同。使用CORP算法,我们预测类似于70%的人类OR基因可能是无功能的假基因,比迄今为止怀疑的要高得多。我们提出的方法也可用于更好地注释其他基因家族中的非活性成员。
Background: Olfactory receptors (ORs), the largest mammalian gene superfamily (900-1400 genes), has > 50% pseudogenes in humans. While most of these inactive genes are identified via coding frame (nonsense) disruptions, seemingly intact genes may also be inactive due to other deleterious (missense) mutations. An ultimate assessment of the actual size of the functional human OR repertoire thus requires an accurate distinction between genes and pseudogenes.Results: To characterize inactive ORs with intact open reading frame, we have developed a probabilistic Classifier for Olfactory Receptor Pseudogenes (CORP). This algorithm is based on deviations from a functionally crucial consensus, constituting sixty highly conserved positions identified by a comparison of two evolutionarily-constrained OR repertoires ( mouse and dog) with a small pseudogene fraction. We used a logistic regression analysis to assign appropriate coefficients to the conserved position and thus achieving maximal separation between active and inactive ORs. Consequently, the algorithms identified only 5% of the mouse functional ORs as pseudogenes, setting an upper limit of 0.05 to the false positive detection. Finally we used this algorithm to classify the 384 purportedly intact human OR genes. Of these, 135 were predicted as likely encoding non-functional proteins, and 38 were segregating between active and inactive forms due to missense polymorphisms.Conclusion: We demonstrated that the CORP algorithm is capable to distinguish between functional and non-functional OR genes with high precision even when the encoded protein would differ by a single amino acid. Using the CORP algorithm, we predict that similar to 70% of human OR genes are likely non-functional pseudogenes, a much higher number than hitherto suspected. The method we present may be employed for better annotation of inactive members in other gene families as well.