课题基金 / 基金详情

BIGDATA: Collaborative Research: F: Efficient and Exact Methods for Big Data Reduction

BIGDATA: Collaborative Research: F: Efficient and Exact Methods for Big Data Reduction
BIGDATA:协作研究:F:大数据缩减的高效且精确的方法
批准号:
1633359
负责人:
Shuiwang Ji
金额:
$45.85万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2019-03-31

项目摘要

项目成果

Shuiwang Ji的其他基金

相似基金

相关文献

中文摘要
翻译
摘要大数据研究涉及分析不断增长的数据集,这些数据集具有大量样本、非常高维的特征向量以及复杂多样的结构。这些数据集不断增长的数量和复杂性使得许多传统技术不足以从中提取知识。一个新兴的领域,称为稀疏学习,通过识别一小部分解释性特征和/或样本,在从大数据中学习方面取得了巨大成功。典型的例子包括选择最能代表用户的功能?推荐系统的偏好,基于成像数据识别预测神经系统疾病的大脑区域,以及从原始图像中提取语义信息用于对象识别。然而,由于稀疏诱导正则化,训练稀疏学习模型可能在计算上是禁止的,这是非平滑的,并且在合并复杂结构时可能非常复杂。该项目旨在开发算法和工具,以显着加快大数据应用稀疏学习模型的训练过程。关键思想是有效地识别冗余特征和/或样本,这些特征和/或样本可以从训练阶段中移除,而不会丢失有用的感兴趣信息。这些独特技术的成功预计将在时间和空间方面以数量级的数量级大幅扩展大数据的稀疏学习。参与者计划将该项目开发的海量数据缩减工具纳入其教育和外联活动,包括开发新课程和将项目组成部分纳入现有课程。本计划的主要技术创新包括:(1)研究员将为输入和输出结构均可用有向无环图表示的一般情况开发有效的特征约简方法;建议的公式包括许多现有的方法作为特殊情况;(2)PI将开发有效的方法,在统一的公式下同时减少特征和样本的数量,也可以包含各种结构;(3)PI将开发有效的方法来丢弃不相关的数据子空间,以加速发现过程大数据中常见的低秩结构。所有提出的数据简化方法都是精确的,即,在简化数据集上学习的模型与在完整数据集上学习的模型相同。该项目在很大程度上依赖于优化理论,特别是灵敏度分析和凸几何。该项目的成果包括一个统一的方法来加速稀疏学习,并为开发高效和精确的数据简化方法提供一个系统的框架。对冗余数据识别的系统研究和深入探索,有望加深对稀疏学习技术的理解,并大大提高其在大数据分析中的应用。
英文摘要
AbstractResearch in big data involves analyzing growing data sets with huge numbers of samples, very high-dimensional feature vectors, and complex and diverse structures. The ever-growing volume and complexity of these data sets make many traditional techniques inadequate to extract knowledge from them. An emerging area, known as sparse learning, has achieved great success in learning from big data by identifying a small set of explanatory features and/or samples. Typical examples include selecting features that are most indicative of users? preferences for recommendation systems, identifying brain regions that are predictive of neurological disorders based on imaging data, and extracting semantic information from raw images for object recognition. However, training sparse learning models can be computationally prohibitive due to the sparsity-inducing regularization, which is non-smooth and can be highly complex when incorporating complex structures. This project aims at developing algorithms and tools to significantly accelerate the training process of sparse learning models for big data applications. The key idea is to efficiently identify redundant features and/or samples, which can be removed from the training phase without losing useful information of interests. Success in these unique techniques is expected to dramatically scaling up sparse learning for big data by orders of magnitude in terms of both time and space. The PIs plan to integrate the big data reduction tools developed in this project into their education and outreach activities, including development of new courses and integration of project components into existing courses. The PIs will make special efforts to recruit female and underrepresented students to this project.The major technical innovations of this project include the following components: (1) the PIs will develop efficient feature reduction methods for the generic scenario where the structures of both input and output can be represented by directed acyclic graphs; the proposed formulations include many existing approaches as special cases; (2) the PIs will develop efficient methods to reduce the numbers of features and samples simultaneously under a unified formulation, which can also incorporate various structures; (3) the PIs will develop efficient methods to discard irrelevant data subspaces to accelerate the process of uncovering low-rank structures commonly seen in big data. All the proposed data reduction methods are exact, i.e., the models learned on the reduced data sets are identical to the ones learned on the full data sets. This project heavily relies on optimization theory, especially on sensitivity analysis and convex geometry. The outcome of this project includes a unified approach to accelerate sparse learning and provide a systematic framework for developing efficient and exact data reduction methods. The systematic study and in-depth exploration of redundant data identification is expected to deepen the understanding of sparse learning techniques and dramatically enhance their applications in big data analytics.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1137/1.9781611975321.67
发表时间: 2017-05
期刊:
影响因子: --
作者: [Zhengyang Wang;Shuiwang Ji]
通讯作者: Zhengyang Wang;Shuiwang Ji
DOI: 10.1145/2939672.2939859
发表时间: 2016-08
期刊: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
影响因子: --
作者: [Qingyang Li;Shuang Qiu;Shuiwang Ji;P. Thompson;Jieping Ye;Jie Wang]
通讯作者: Qingyang Li;Shuang Qiu;Shuiwang Ji;P. Thompson;Jieping Ye;Jie Wang
III: Small: 3D Graph Neural Networks: Completeness, Efficiency, and Applications
Collaborative Research: ABI Innovation: Towards Computational Exploration of Large-Scale Neuro-Morphological Datasets
III: Small: Collaborative Research: Demystifying Deep Learning on Graphs: From Basic Operations to Applications
III: Medium: Collaborative Research: Towards Scalable and Interpretable Graph Neural Networks
海外基金