课题基金 / 基金详情

CIF: CAREER: Robust, Interpretable, and Efficient Unsupervised Learning with K-set Clustering

CIF: CAREER: Robust, Interpretable, and Efficient Unsupervised Learning with K-set Clustering
CIF:职业:使用 K 集聚类进行稳健、可解释且高效的无监督学习
批准号:
1845076
负责人:
Laura Balzano
金额:
$59.68万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2019
资助国家:
美国
项目状态:
未结题
起止时间:
2019-05-01 至 2025-04-30

项目摘要

项目成果

Laura Balzano的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Modern machine learning techniques aim to design models and algorithms that allow computers to learn efficiently from vast amounts of previously unexplored data. These problems are called 'unsupervised' because no human-provided information about the data is available to guide the machine learning process. Arguably the two most important unsupervised machine learning tools are dimensionality-reduction and clustering. In dimensionality-reduction, the algorithm seeks a simple low-dimensional structure that captures the interesting behavior in the data. In clustering, the algorithm seeks to group data points together into meaningful clusters. As increasingly higher-dimensional data are collected about progressively more elaborate physical, biological, and social phenomena, algorithms that aim at both dimensionality reduction and clustering are often highly applicable. However, joint formulations in the literature are often ad-hoc and fundamentally unable to operate on real data that have missing elements, corruptions, and heterogeneity --- critical machine learning challenges for modern data problems. This research project is expected to have broad applicability in data science, and will be demonstrated in two applications: genetics and computer vision. The joint clustering and dimensionality reduction formulation used in this project, called K-set clustering, seeks K "central sets" constrained to have some low-dimensional representation, each of which represents one of K clusters in the data. The formulation is a generalization of K-means, K-subspaces, and principal component analysis, and it naturally leads to several novel problem instances. Given a defined set geometry, the corresponding problem instance is approached from two perspectives: understanding the geometry of that instance of the problem formulation, and learning those geometric models from data. Three specific examples of the problem formulation will be studied: subspace clustering, variety clustering, and polyhedral set clustering. While each example presents intrinsic and unique challenges, these are just examples of a larger paradigm that is limited only by one's ability to define sets amenable to modeling the geometric structure in data. The formulation allows for interpretable data analysis, with a framework that can readily incorporate missing data and heterogeneous data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(25)
专著(0)
科研奖励(0)
会议论文
DOI: 10.48550/arxiv.2209.09211
发表时间: 2022-09
期刊: ArXiv
影响因子: --
作者: [Can Yaras;Peng Wang;Zhihui Zhu;L. Balzano;Qing Qu]
通讯作者: Can Yaras;Peng Wang;Zhihui Zhu;L. Balzano;Qing Qu
DOI: 10.48550/arxiv.2205.02215
发表时间: 2022-05
期刊:
影响因子: --
作者: [Davoud Ataee Tarzanagh;Mingchen Li;Christos Thrampoulidis;Samet Oymak]
通讯作者: Davoud Ataee Tarzanagh;Mingchen Li;Christos Thrampoulidis;Samet Oymak
DOI: 10.1109/jproc.2020.3021381
发表时间: 2020-05
期刊: Proceedings of the IEEE
影响因子: 20.6
作者: [M. Nokleby;Haroon Raja;W. Bajwa]
通讯作者: M. Nokleby;Haroon Raja;W. Bajwa
DOI: --
发表时间: 2020-02
期刊:
影响因子: --
作者: [Amanda Bower;L. Balzano]
通讯作者: Amanda Bower;L. Balzano
20
    CIF: Small: Learning Low-Dimensional Representations with Heteroscedastic Data Sources
    BRIGE: Simultaneous Modeling and Calibration for Environmental Sensor Data
    海外基金