Adaptive Thresholding for Hierarchical Clustering of Variables, with Connections to Scan Statistics
Adaptive Thresholding for Hierarchical Clustering of Variables, with Connections to Scan Statistics
批准号:
1613202
负责人:
Max G'Sell
金额:
$15.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-08-01 至 2020-07-31
中文摘要
在使用大数据集的现代数据分析中,一个常见的目标是检测表现出相似行为的变量组。此任务通常称为集群。例如,在遗传学和蛋白质组学中,聚类可以揭示科学上感兴趣的结构,例如潜在的生物路径。除了检测数据中科学相关的结构外,聚类还可以用于简化数据表示和分析。最广泛使用的聚类方法之一称为层次聚类。在层次聚类中,计算每对变量之间的相似性度量,如相关性,然后重复合并相似的变量组。这就引出了一个基本问题:应该进行多少分组?拟议的研究包括几个项目,旨在根据数据中存在的相似程度,制定广泛适用的方法,以确定适当的聚类量。由此产生的程序还将为结果组的意义提供统计保证。这项建议旨在开发实用的程序,当应用于变量的成对相似性时,对层次聚类树形图进行自适应阈值处理。这些过程将与关于所得到的聚类的错误聚类错误率的推论保证相关联。结果将针对一系列共同联系和可变相似性衡量标准。PI还将在现代遗传学应用中演示这些过程。为了支持这些过程,将发展新的理论来描述变量相似性度量的大序统计量,包括它们联合分布的新的渐近界和它们的最大值的新的有限样本界。这里提出的技术也将适用于统计学中的其他基于阈值的程序;特别是,可以在拟议的工作和扫描统计的自适应阈值程序之间建立联系。
英文摘要
In modern data analysis with large data sets, a common goal is to detect groups of variables that exhibit similar behavior. This task is usually referred to as clustering. In genetics and proteomics, for instance, clustering can reveal structures of scientific interest, such as potential biological pathways. On top of detecting scientifically relevant structure in the data, clustering can also be used to simplify data representations and analysis. One of the most widely used approaches to clustering is called hierarchical clustering. In hierarchical clustering, a measure of similarity, like correlation, is computed between each pair of variables, and then similar groups of variables are repeatedly merged. This leads to a fundamental question: how much grouping should be done? The proposed research consists of several projects aimed at developing broadly applicable methods for determining the appropriate amount of clustering, based on the degree of similarity present in the data. The resulting procedures will also provide statistical guarantees on the meaning of the resulting groups.This proposal aims to develop practical procedures for adaptive thresholding of hierarchical clustering dendrograms, when applied to pairwise similarities of variables. These procedures will be connected to inferential guarantees about the false cluster error rate of the resulting clustering. The results will target a range of common linkages and variable similarity measures. The PI will also demonstrate these procedures in a modern genetics application.To support these procedures, new theory will be developed describing the large order statistics of variable similarity measures, including new asymptotic bounds on their joint distributions and new finite-sample bounds on their maxima. The techniques proposed here will also have application to other threshold-based procedures in statistics; in particular, connections may be made between the proposed work and adaptive thresholding procedures for scan statistics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
NCS-FO:Collaborative Research:Decoding and Reconstructing the Neural Basis of Real World Social Perception
-
批准号:1734868
-
项目类别:Standard Grant
-
资助金额:$49.01万
-
财政年份:2017
-
负责人:Max G'Sell
-
依托单位:
海外基金