课题基金 / 基金详情

High-dimensional Clustering: Theory and Methods

High-dimensional Clustering: Theory and Methods
高维聚类:理论与方法
批准号:
1713003
负责人:
Sivaraman Balakrishnan
金额:
$38.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-07-01 至 2021-06-30

项目摘要

项目成果

Sivaraman Balakrishnan的其他基金

相似基金

相关文献

中文摘要
翻译
在过去的二十年里,科学和工程领域出现了数据集规模和复杂性的爆炸式增长。从广义上讲,发现数据中潜在结构的聚类方法是我们导航,探索和可视化海量数据集的主要工具。这些方法已广泛而成功地应用于遗传学、医学、精神病学、考古学和人类学、植物社会学、经济学等领域。尽管它无处不在,聚类方法的广泛科学采用受到了阻碍,缺乏灵活的聚类方法,高维数据集和聚类问题缺乏有意义的推理保证。因此,本研究的目标是开发新的和有效的方法来聚类复杂的数据集,并进一步发展推理的基础-这反过来又会导致可操作的结论-这些方法。这项研究将导致新的聚类方法的发展,以及更深入地了解旨在揭示数据中潜在结构的方法的基本局限性。该项目的研究部分包括四个目标,旨在解决这一高级目标的相关方面:(a)分析和开发新的高维数据集聚类方法,特别关注实际有用的方法,如基于混合模型的聚类和最小体积聚类;(B)在聚类的背景下开发新的推理方法,受科学应用的推动,在这些应用中,不仅要对数据进行聚类,而且要清楚地描述所发现的聚类的采样可变性;(c)开发高维聚类的基本下界(d)开发具有推理保证的功能数据聚类的新方法。这些研究组成部分与具体的教育举措密切相关,包括开发和广泛传播用于高维聚类的公开软件;机器学习会议上的教程和研讨会,以及促进卡内基梅隆大学统计和机器学习部门之间的进一步互动。
英文摘要
The past two decades have witnessed an explosion in the scale and complexity of data sets that arise in science and engineering. Broadly, clustering methods which discover latent structure in data are our primary tool for navigating, exploring and visualizing massive datasets. These methods have been widely and successfully applied in phylogeny, medicine, psychiatry, archaeology and anthropology, phytosociology, economics and several other fields. Despite its ubiquity, the widespread scientific adoption of clustering methods have been hindered by the lack of flexible clustering methods for high-dimensional datasets and by the dearth of meaningful inferential guarantees in clustering problems. Accordingly, the goal of this research is to develop new and effective methods for clustering complex data-sets, and to further develop an inferential grounding -- which will in turn lead to actionable conclusions -- for these methods. This research will lead to the development of new clustering methods, as well as to a deeper understanding of the fundamental limitations of methods aimed at uncovering latent structure in data. The research component of this project consists of four aims designed to address related aspects of this high-level goal: (a) analyze and develop new clustering methods for high-dimensional datasets, with a particular focus on practically useful methods like mixture-model based clustering, and minimum volume clustering; (b) develop novel methods for inference in the context of clustering, motivated by scientific applications where it is important not only to cluster the data but also to clearly characterize the sampling variability of the discovered clusters; (c) develop fundamental lower bounds for high-dimensional clustering (d) develop novel methods for clustering functional data with inferential guarantees. These research components are closely coupled with concrete educational initiatives, including the development and broad dissemination of publicly-available software for high-dimensional clustering; tutorials and workshops at Machine Learning conferences and fostering further interactions between the Departments of Statistics and Machine Learning at Carnegie Mellon.
期刊论文(24)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1214/19-ejs1639
发表时间: 2018-12
期刊: Electronic Journal of Statistics
影响因子: 1.1
作者: [I. Verdinelli;L. Wasserman]
通讯作者: I. Verdinelli;L. Wasserman
DOI: 10.1214/18-ejs1510
发表时间: 2018-05
期刊: Electronic Journal of Statistics
影响因子: 1.1
作者: [I. Verdinelli;L. Wasserman]
通讯作者: I. Verdinelli;L. Wasserman
DOI: 10.1214/20-aos2030
发表时间: 2020-01
期刊: The Annals of Statistics
影响因子: --
作者: [Matey Neykov;Sivaraman Balakrishnan;L. Wasserman]
通讯作者: Matey Neykov;Sivaraman Balakrishnan;L. Wasserman
DOI: 10.1016/j.jmva.2019.06.004
发表时间: 2017-02
期刊: J. Multivar. Anal.
影响因子: --
作者: [Yining Wang;Jialei Wang;Sivaraman Balakrishnan;Aarti Singh]
通讯作者: Yining Wang;Jialei Wang;Sivaraman Balakrishnan;Aarti Singh
共 20 条
    Foundations of High-Dimensional and Nonparametric Hypothesis Testing
    • 批准号:
      2113684
    • 项目类别:
      Standard Grant
    • 资助金额:
      $25.0万
    • 财政年份:
      2021
    • 负责人:
      Sivaraman Balakrishnan
    • 依托单位:
    海外基金