课题基金 / 基金详情

BIGDATA: Collaborative Research: F: Efficient and Exact Methods for Big Data Reduction

BIGDATA: Collaborative Research: F: Efficient and Exact Methods for Big Data Reduction
BIGDATA:协作研究:F:大数据缩减的高效且精确的方法
批准号:
1633370
负责人:
Qiaozhu Mei
金额:
$49.97万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2022-08-31

项目摘要

项目成果

Qiaozhu Mei的其他基金

相似基金

相关文献

中文摘要
翻译
大数据的研究涉及分析不断增长的数据集,这些数据集具有大量的样本、非常高维的特征向量以及复杂多样的结构。这些数据集的数量和复杂性不断增长,使得许多传统技术无法从其中提取知识。一个名为稀疏学习的新兴领域通过识别一小组解释性特征和/或样本,在从大数据中学习方面取得了巨大成功。典型的例子包括选择最能代表用户的功能?对推荐系统的偏好,基于成像数据识别预测神经疾病的大脑区域,以及从原始图像中提取语义信息用于对象识别。然而,由于稀疏诱导正则化,训练稀疏学习模型可能在计算上被禁止,这是非光滑的,并且在结合复杂结构时可能非常复杂。该项目旨在开发算法和工具,以显著加快用于大数据应用的稀疏学习模型的训练过程。其关键思想是有效地识别冗余特征和/或样本,这些特征和/或样本可以在不丢失感兴趣的有用信息的情况下从训练阶段移除。这些独特技术的成功预计将极大地扩大大数据稀疏学习在时间和空间上的数量级。投资促进机构计划将在该项目中开发的大数据减少工具纳入其教育和外联活动,包括开发新课程和将项目组成部分纳入现有课程。这个项目的主要技术创新包括以下几个部分:(1)个人信息系统将为输入和输出的结构都可以用有向无环图表示的通用场景开发有效的特征约简方法;建议的公式包括许多现有的方法作为特例;(2)个人信息系统将开发有效的方法,在统一的公式下同时减少特征和样本的数量,该方法也可以包含各种结构;(3)PI将开发有效的方法来丢弃不相关的数据子空间,以加快发现大数据中常见的低等级结构的进程。所有提出的数据约简方法都是精确的,即在约简的数据集上学习的模型与在全数据集上学习的模型是相同的。这个项目在很大程度上依赖于最优化理论,特别是灵敏度分析和凸几何。该项目的成果包括一种加速稀疏学习的统一方法,并为开发高效和准确的数据简化方法提供了一个系统框架。对冗余数据识别的系统研究和深入探索,有望加深对稀疏学习技术的理解,大幅提升稀疏学习技术在大数据分析中的应用。
英文摘要
Research in big data involves analyzing growing data sets with huge numbers of samples, very high-dimensional feature vectors, and complex and diverse structures. The ever-growing volume and complexity of these data sets make many traditional techniques inadequate to extract knowledge from them. An emerging area, known as sparse learning, has achieved great success in learning from big data by identifying a small set of explanatory features and/or samples. Typical examples include selecting features that are most indicative of users? preferences for recommendation systems, identifying brain regions that are predictive of neurological disorders based on imaging data, and extracting semantic information from raw images for object recognition. However, training sparse learning models can be computationally prohibitive due to the sparsity-inducing regularization, which is non-smooth and can be highly complex when incorporating complex structures. This project aims at developing algorithms and tools to significantly accelerate the training process of sparse learning models for big data applications. The key idea is to efficiently identify redundant features and/or samples, which can be removed from the training phase without losing useful information of interests. Success in these unique techniques is expected to dramatically scaling up sparse learning for big data by orders of magnitude in terms of both time and space. The PIs plan to integrate the big data reduction tools developed in this project into their education and outreach activities, including development of new courses and integration of project components into existing courses. The PIs will make special efforts to recruit female and underrepresented students to this project.The major technical innovations of this project include the following components: (1) the PIs will develop efficient feature reduction methods for the generic scenario where the structures of both input and output can be represented by directed acyclic graphs; the proposed formulations include many existing approaches as special cases; (2) the PIs will develop efficient methods to reduce the numbers of features and samples simultaneously under a unified formulation, which can also incorporate various structures; (3) the PIs will develop efficient methods to discard irrelevant data subspaces to accelerate the process of uncovering low-rank structures commonly seen in big data. All the proposed data reduction methods are exact, i.e., the models learned on the reduced data sets are identical to the ones learned on the full data sets. This project heavily relies on optimization theory, especially on sensitivity analysis and convex geometry. The outcome of this project includes a unified approach to accelerate sparse learning and provide a systematic framework for developing efficient and exact data reduction methods. The systematic study and in-depth exploration of redundant data identification is expected to deepen the understanding of sparse learning techniques and dramatically enhance their applications in big data analytics.
期刊论文(30)
专著(0)
科研奖励(0)
会议论文
Explainable Prediction of Text Complexity: The Missing Preliminaries for Text Simplification
文本复杂性的可解释预测:文本简化缺失的预备知识
DOI: 10.18653/v1/2021.acl-long.88
发表时间: 2021
期刊: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing
影响因子: --
作者: [Garbacea, Cristina, Guo, Mengtian, Carton, Samuel, Mei, Qiaozhu]
通讯作者: Mei, Qiaozhu
DOI: 10.1145/3425866
发表时间: 2021-03
期刊: ACM Transactions on Internet Technology
影响因子: 5.3
作者: [LiuXuanzhe;WangShangguang;MaYun;Zhangying;MeiQiaozhu;LiuYunxin;Huanggang]
通讯作者: LiuXuanzhe;WangShangguang;MaYun;Zhangying;MeiQiaozhu;LiuYunxin;Huanggang
DOI: 10.18653/v1/d18-1386
发表时间: 2018-09
期刊: ArXiv
影响因子: --
作者: [Samuel Carton;Qiaozhu Mei;P. Resnick]
通讯作者: Samuel Carton;Qiaozhu Mei;P. Resnick
DOI: --
发表时间: 2021-06
期刊: Mathematical Programming
影响因子: 2.7
作者: [Jiaqi Ma;Junwei Deng;Qiaozhu Mei]
通讯作者: Jiaqi Ma;Junwei Deng;Qiaozhu Mei
23
    NSF Student Travel Grant for 2017 Conference on Knowledge Discovery and Data Mining (KDD 2017)
    CiC (RDDC): Wordsmith in the Cloud - Refining Language Models Using Web-Scale Language Networks
    CAREER: Eyes of the Foreseer - Integrative and In Situ Information Retrieval and Mining in Online Communities
    SoCS: Assessing Information Credibility Without Authoritative Sources
    海外基金