Statistical significance for hierarchical clustering in genetic association and microarray expression studies.

Statistical significance for hierarchical clustering in genetic association and microarray expression studies.
复制标题

DOI:
10.1186/1471-2105-4-62
复制
发表时间:
2003-12-11
期刊:
影响因子:
3
通讯作者:
Ott J
Ott J
中科院分区:
生物学4区
文献类型:
--
作者:
Levenstien MA;Yang Y;Ott J

文献摘要

参考文献

被引文献

相似文献

随着分子遗传学实验室产生的数据量越来越多,由于研究的结果或变量数量庞大,因此通常很难理解结果。实例包括在大量基因座处的大量基因和单倍型的表达水平。然后,自然地将观察分组为更小数量的类,以便更容易地概述和解释数据。这种分组通常在多个步骤中进行,并借助分层聚类分析,每个步骤通过组合类似的观察或类来产生较少数量的类。在每一步中,无论是隐式的还是显式的,研究人员倾向于解释结果,并最终集中在提供“最佳”(最显著)结果的类集合上。虽然这种方法是有意义的,但实验的总体统计意义必须包括聚类过程,该过程修改了数据的分组结构,并经常消除变异。对于分层聚类的数据,我们建议考虑最强的结果,或者等效地,最小的p值作为感兴趣的实验统计量,并评估其显著性水平,以进行统计显著性的全局评估。我们将我们的方法应用于单倍型关联和微阵列表达研究的数据集,其中已使用分层聚类。在我们研究的所有情况下,我们发现,依赖于一组类在聚类过程中导致显着性水平太小,当与整体统计相关联的显着性水平,结合了聚类过程相比。换句话说,依赖于一个聚类步骤可能会提供一个形式上有意义的结果,而整个实验并不显着。
With the increasing amount of data generated in molecular genetics laboratories, it is often difficult to make sense of results because of the vast number of different outcomes or variables studied. Examples include expression levels for large numbers of genes and haplotypes at large numbers of loci. It is then natural to group observations into smaller numbers of classes that allow for an easier overview and interpretation of the data. This grouping is often carried out in multiple steps with the aid of hierarchical cluster analysis, each step leading to a smaller number of classes by combining similar observations or classes. At each step, either implicitly or explicitly, researchers tend to interpret results and eventually focus on that set of classes providing the "best" (most significant) result. While this approach makes sense, the overall statistical significance of the experiment must include the clustering process, which modifies the grouping structure of the data and often removes variation. For hierarchically clustered data, we propose considering the strongest result or, equivalently, the smallest p-value as the experiment-wise statistic of interest and evaluating its significance level for a global assessment of statistical significance. We apply our approach to datasets from haplotype association and microarray expression studies where hierarchical clustering has been used. In all of the cases we examine, we find that relying on one set of classes in the course of clustering leads to significance levels that are too small when compared with the significance level associated with an overall statistic that incorporates the process of clustering. In other words, relying on one step of clustering may furnish a formally significant result while the overall experiment is not significant.
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1038/35000501
发表时间: 2000-02-03
期刊: NATURE
影响因子: 64.8
作者:
Alizadeh, AA;Eisen, MB;Staudt, LM
通讯作者: Staudt, LM
DOI: 10.1101/gr.204001
发表时间: 2001-12-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Hoh, J;Wille, A;Ott, J
通讯作者: Ott, J
DOI: 10.1073/pnas.241500798
发表时间: 2001-11-20
影响因子: 11.1
作者:
Garber, ME;Troyanskaya, OG;Petersen, I
通讯作者: Petersen, I
DOI: 10.1093/bioinformatics/btf877
发表时间: 2003-02-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Reiner, A;Yekutieli, D;Benjamini, Y
通讯作者: Benjamini, Y