Empirical study of supervised gene screening

Empirical study of supervised gene screening
复制标题

DOI:
10.1186/1471-2105-7-537
复制
发表时间:
2006-12-18
期刊:
影响因子:
3
通讯作者:
Ma, Shuangge
Ma, Shuangge
中科院分区:
生物学4区
文献类型:
--
作者:
Ma, Shuangge

文献摘要

被引文献

相似文献

背景:微阵列研究提供了一种将表型变异与其遗传原因联系起来的方法。使用高维微阵列测量构建预测模型通常包括三个步骤:(1)无监督基因筛选;(2)监督基因筛选;和(3)统计模型构建。基于边缘基因排序的监督基因筛选是一种常用的减少模型构建中基因数量的方法。各种简单的统计量,如t-统计量或信噪比,已被用来在监督筛选基因排序。尽管监督基因筛选的应用非常广泛,但其统计学研究仍然很少。我们的研究部分是出于使用不同的监督基因screeningmethods.Results所造成的基因发现结果的差异:我们调查的一致性和重复性的监督基因screening基于8个常用的边缘统计。一致性通过使用不同边际统计筛选的排名靠前的基因之间的重叠的相对分数来评估。我们提出了一个Bootstrap复制指数,它衡量的监督筛选下的单个基因的再现性。实证研究是基于四个公开的微阵列数据。我们考虑的情况下,前20%,40%和60%的基因screen.Conclusion:从基因发现的角度来看,监督基因筛选的效果,基于不同的边缘统计是不可忽视的。实证研究表明:(1)通过不同监督筛选的基因可能有很大差异;(2)一致性可能会有所不同,这取决于基础数据结构和所选基因的百分比;(3)用Bootstrap Reproductivity Index进行评估,通过监督筛选的基因仅具有中等可重复性;(4)基于可重复性的监督筛选无法改善一致性。
Background: Microarray studies provide a way of linking variations of phenotypes with their genetic causations. Constructing predictive models using high dimensional microarray measurements usually consists of three steps: (1) unsupervised gene screening; (2) supervised gene screening; and (3) statistical model building. Supervised gene screening based on marginal gene ranking is commonly used to reduce the number of genes in the model building. Various simple statistics, such as t-statistic or signal to noise ratio, have been used to rank genes in the supervised screening. Despite of its extensive usage, statistical study of supervised gene screening remains scarce. Our study is partly motivated by the differences in gene discovery results caused by using different supervised gene screening methods.Results: We investigate concordance and reproducibility of supervised gene screening based on eight commonly used marginal statistics. Concordance is assessed by the relative fractions of overlaps between top ranked genes screened using different marginal statistics. We propose a Bootstrap Reproducibility Index, which measures reproducibility of individual genes under the supervised screening. Empirical studies are based on four public microarray data. We consider the cases where the top 20%, 40% and 60% genes are screened.Conclusion: From a gene discovery point of view, the effect of supervised gene screening based on different marginal statistics cannot be ignored. Empirical studies show that (1) genes passed different supervised screenings may be considerably different; (2) concordance may vary, depending on the underlying data structure and percentage of selected genes; (3) evaluated with the Bootstrap Reproducibility Index, genes passed supervised screenings are only moderately reproducible; and ( 4) concordance cannot be improved by supervised screening based on reproducibility.