课题基金 / 基金详情

RI: Medium: Extreme Clustering

RI: Medium: Extreme Clustering
RI:中:极端集群
批准号:
1763618
负责人:
Andrew McCallum
金额:
$110.39万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-09-01 至 2023-08-31

项目摘要

项目成果

Andrew McCallum的其他基金

相似基金

相关文献

中文摘要
翻译
聚类是数据科学、假设发现、模式发现和信息集成的基本工具。给定一个对象集合,集群是自动对对象进行分组的任务,以便组(称为集群)中的对象彼此之间的相似性比其他集群中的对象更相似。聚类在医学、工程、科学和商业等领域有着广泛的应用。大多数现代聚类方法可以很好地扩展到大量对象,但不能扩展到大量集群。此外,目前广泛使用的聚类方法将对象过多地分配到聚类中,精度不高,并且不代表聚类中的不确定性。所有这些弱点限制了许多科学、工程和其他高影响应用程序的分析能力。这个项目正在为大规模集群开发新的机器学习和算法,可以扩展到大量的对象和大量的集群。该项目将建立在最近初步成功的一系列算法的基础上,这些算法构建分层聚类,支持有效地将数据重新分配到新的聚类,并且自然地代表不确定性。这项新研究旨在进一步提高准确性和可扩展性。项目团队将展示其在与国家优先事项相关的多个领域的新研究,包括用于材料科学发现的化合物聚类、单细胞基因组数据聚类以及科学元数据(如论文作者、专利作者、论文、机构等)的实体解析——创建促进科学发现、合作和科学同行评审的工具。作为该项目的一部分开发的所有软件都将作为开源软件发布,以便在研究和实践中促进实验和采用我们的方法。项目团队将开发一个机器学习和算法交叉的教程,并将额外向计算机科学家以外的研究人员教授一门关于高效聚类方法的课程。该项目将开发机器学习和分层聚类算法的新研究,该算法可扩展到大量输入对象N和大量聚类K,这是一个被称为“极端聚类”的问题设置,以其类似的动机监督表表“极端分类”命名。该项目建立在最近对PERCH初步工作的成功基础上,这是一组用于大规模、增量数据、非贪婪、分层聚类的算法,已经取得了令人瞩目的最新成果。该方法有效地将新数据点路由到增量构建树的叶子。出于对准确性和速度的要求,该方法执行树旋转,以提高子树的纯度,并鼓励平衡树。实验表明,与其他构建树的聚类算法相比,该算法构建的树更精确,并且在N和K下都具有良好的可扩展性,在近一半的时间内实现了比最强的聚类竞争对手更高的聚类质量。该项目将进行新的研究(a)通过替代聚类成本函数和数据表示来提高灵活性,(b)通过新的树路由函数进一步提高可扩展性和准确性,(c)开发新的树切方法来确定最佳聚类和聚类分布,以及(d)发明多种相互关联的数据实例类型的联合聚类的新方法。该研究的评估和应用将在多个广泛影响的大数据领域进行,包括生物医学、材料科学、图像分析和科学信息集成。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Clustering is a fundamental tool for data science, hypothesis discovery, pattern discovery, and information integration. Given a collection of objects, clustering is the task of automatically grouping the objects so that objects within a group (called a cluster) are more similar to each other than to objects in other clusters. Clustering is widely used in medicine, engineering, science and commerce. Most modern clustering methods scale well to a large number of objects, but not to a large number of clusters. Furthermore currently widely used clustering methods excessively assign objects to clusters, suffer in accuracy, and do not represent uncertainty in the clustering. All of these weaknesses limit analysis capabilities in many scientific, engineering, and other high-impact applications. This project is developing new machine learning and algorithms for large-scale clustering that scales to both massive number of objects and massive number of clusters. This project will build on the recent preliminary success with a family of algorithms that build hierarchical clustering, which supports efficient re-assignment of data to new clusters, and which naturally represents uncertainty. The new research aims to further increase accuracy and scalability. The project team will demonstrate its new research in multiple domains relevant to national priorities, including clustering chemical compounds for material science discovery, clustering single cell genome data, and entity resolution on scientific metadata (such as paper authors, patent authors, papers, institutions, etc)--- creating tools that advance scientific discovery, collaboration and scientific peer review. All of the software developed as part of this project will be released as open source software in order to facilitate experimentation and adoption of our methods in research and practice. The project team will develop a tutorial at the intersection of machine learning and algorithms, and will additionally teach a course on efficient clustering methods to researchers beyond computer scientists.This project will develop new research on machine learning and algorithms for hierarchical clustering that scales to both massive number of input objects, N, and massive number of clusters, K---a problem setting termed "extreme clustering," named after its similarly- motivated supervised cousin, "extreme classification." The project builds on the successes of recent preliminary work on PERCH, a family of algorithms for large-scale, incremental-data, non-greedy, hierarchical clustering that has achieved remarkable new state-of-the-art results. The method efficiently routes new data points to the leaves of an incrementally-built tree. Motivated by the desire for both accuracy and speed, the approach performs tree rotations both for the sake of enhancing subtree purity and encouraging balanced trees. Experiments demonstrate that PERCH constructs more accurate trees than other tree-building clustering algorithms and scales well with both N and K, achieving a higher quality clustering than the strongest at clustering competitor in nearly half the time. The project will perform new research (a) improving flexibility through alternative clustering cost functions and data representations, (b) further improving scalability and accuracy through new tree routing functions, (c) developing new tree-cut methods for determining the best clusterings and distributions over clusterings, and (d) inventing new methods for joint clustering of multiple inter-related data instance types. Evaluation and application of the research will be conducted on multiple broad-impact, large-data domains, including biomedicine, material science, image analysis, and scientific information integration.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(23)
专著(0)
科研奖励(0)
会议论文
Exact and Approximate Hierarchical Clustering with A*
使用 A* 的精确和近似层次聚类
DOI: --
发表时间: 2021
期刊: The Conference on Uncertainty in Artificial Intelligence (UAI
影响因子: --
作者: [Greenberg, Craig, Macaluso, Sebastian, Monath, Nicholas, Dubey, Avinava, Flaherty, Patrick, Zaheer, Manzil, Ahmed, Amr, Cranmer, Kyle, McCallum, Andrew]
通讯作者: McCallum, Andrew
DOI: --
发表时间: 2021
期刊:
影响因子: --
作者: [Nicholas Monath;M. Zaheer;Kumar Avinava Dubey;Amr Ahmed;A. McCallum]
通讯作者: Nicholas Monath;M. Zaheer;Kumar Avinava Dubey;Amr Ahmed;A. McCallum
DOI: 10.1145/3292500.3330929
发表时间: 2019-07
期刊: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
影响因子: --
作者: [Nicholas Monath;Ari Kobren;A. Krishnamurthy;Michael R. Glass;A. McCallum]
通讯作者: Nicholas Monath;Ari Kobren;A. Krishnamurthy;Michael R. Glass;A. McCallum
DOI: --
发表时间: 2022
期刊:
影响因子: --
作者: [Dongxu Zhang;Michael Boratko;Cameron Musco;A. McCallum]
通讯作者: Dongxu Zhang;Michael Boratko;Cameron Musco;A. McCallum
共 21 条
    Collaborative Research: SOS-DCI / HNDS-R: Advancing Semantic Network Analysis to Better Understand How Evaluative Exchanges Shape Scientific Arguments
    • 批准号:
      2244805
    • 项目类别:
      Standard Grant
    • 资助金额:
      $22.5万
    • 财政年份:
      2023
    • 负责人:
      Andrew McCallum
    • 依托单位:
    RI: Medium: Probabilistic Box Embeddings
    • 批准号:
      2106391
    • 项目类别:
      Standard Grant
    • 资助金额:
      $84.99万
    • 财政年份:
      2021
    • 负责人:
      Andrew McCallum
    • 依托单位:
    DMREF: Collaborative Research: The Synthesis Genome: Data Mining for Synthesis of New Materials
    • 批准号:
      1922090
    • 项目类别:
      Standard Grant
    • 资助金额:
      $40.0万
    • 财政年份:
      2019
    • 负责人:
      Andrew McCallum
    • 依托单位:
    DMREF: Collaborative Research: The Synthesis Genome: Data Mining for Synthesis of New Materials
    • 批准号:
      1534431
    • 项目类别:
      Standard Grant
    • 资助金额:
      $36.39万
    • 财政年份:
      2015
    • 负责人:
      Andrew McCallum
    • 依托单位:
    海外基金