Statistics for approximate gene clusters.

Statistics for approximate gene clusters.
复制标题

DOI:
10.1186/1471-2105-14-s15-s14
复制
发表时间:
2013
期刊:
影响因子:
3
通讯作者:
Böcker S
Böcker S
中科院分区:
生物学4区
文献类型:
--
作者:
Jahn K;Winter S;Stoye J;Böcker S

文献摘要

被引文献

相似文献

共定位于多个基因组中的基因可以作为基因组组织的功能限制或残余祖先基因顺序的有力指标。这些模式(通常称为基因簇)的计算检测在过去十年中变得越来越敏感。最强大的方法允许各种类型的不完美簇保护:簇位置可以在内部重新排列。各个簇位置可能仅包含簇基因的子集,并且可能被不涉及的基因破坏。此外,簇位置可能根本不出现在一些甚至大多数研究的基因组中。对这种低质量簇的检测增加了将偶然出现的微弱模式误认为真实发现的风险。因此,评估计算基因簇预测的重要性并区分真正的保守聚类和巧合聚类至关重要。在本文中,我们提出了一种有效且准确的方法来估计近似公共区间模型下基因簇预测的显着性。给定一个单基因簇预测,我们计算在随机基因顺序的零假设下以相同或更高程度的保守性观察到它的概率,并添加校正因子以考虑多重测试。我们的方法考虑了定义基因簇保守质量的所有参数:簇出现的基因组数量、涉及的基因数量、不同基因组中的保守程度以及每个基因组内聚集基因的频率。我们应用我们的方法来评估大量注释良好的基因组中的基因簇预测。
Genes occurring co-localized in multiple genomes can be strong indicators for either functional constraints on the genome organization or remnant ancestral gene order. The computational detection of these patterns, which are usually referred to as gene clusters, has become increasingly sensitive over the past decade. The most powerful approaches allow for various types of imperfect cluster conservation: Cluster locations may be internally rearranged. The individual cluster locations may contain only a subset of the cluster genes and may be disrupted by uninvolved genes. Moreover cluster locations may not at all occur in some or even most of the studied genomes. The detection of such low quality clusters increases the risk of mistaking faint patterns that occur merely by chance for genuine findings. Therefore, it is crucial to estimate the significance of computational gene cluster predictions and discriminate between true conservation and coincidental clustering. In this paper, we present an efficient and accurate approach to estimate the significance of gene cluster predictions under the approximate common intervals model. Given a single gene cluster prediction, we calculate the probability to observe it with the same or a higher degree of conservation under the null hypothesis of random gene order, and add a correction factor to account for multiple testing. Our approach considers all parameters that define the quality of gene cluster conservation: the number of genomes in which the cluster occurs, the number of involved genes, the degree of conservation in the different genomes, as well as the frequency of the clustered genes within each genome. We apply our approach to evaluate gene cluster predictions in a large set of well annotated genomes.