BIGDATA: Collaborative Research: F: Efficient and Exact Methods for Big Data Reduction
BIGDATA: Collaborative Research: F: Efficient and Exact Methods for Big Data Reduction
批准号:
1633370
负责人:
Qiaozhu Mei
金额:
$49.97万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2022-08-31
中文摘要
大数据的研究涉及分析不断增长的数据集,这些数据集具有大量的样本、非常高维的特征向量和复杂多样的结构。这些数据集的数量和复杂性不断增长,使得许多传统技术无法从中提取知识。一个新兴的领域,被称为稀疏学习,通过识别一小部分解释特征和/或样本,在从大数据中学习方面取得了巨大的成功。典型的例子包括选择最能指示用户的功能。推荐系统的偏好,基于成像数据识别预测神经系统疾病的大脑区域,以及从原始图像中提取语义信息用于对象识别。然而,训练稀疏学习模型可能在计算上令人望而却步,因为稀疏性诱导的正则化是不光滑的,并且在合并复杂结构时可能非常复杂。该项目旨在开发算法和工具,以显著加快大数据应用中稀疏学习模型的训练过程。关键思想是有效地识别冗余的特征和/或样本,这些特征和/或样本可以在不丢失有用的兴趣信息的情况下从训练阶段删除。这些独特技术的成功有望在时间和空间方面显著扩大大数据稀疏学习的数量级。项目负责人计划将该项目开发的大数据缩减工具整合到他们的教育和推广活动中,包括开发新课程和将项目组成部分整合到现有课程中。学院将特别努力招收女性和代表性不足的学生参加该项目。本项目的主要技术创新包括以下组成部分:(1)pi将为输入和输出的结构都可以用有向无环图表示的通用场景开发有效的特征约简方法;拟议的提法包括许多作为特殊情况的现有办法;(2) pi将开发在统一公式下同时减少特征和样本数量的有效方法,该方法也可以包含各种结构;(3) pi将开发有效的方法来丢弃不相关的数据子空间,以加速发现大数据中常见的低秩结构的过程。所有提出的数据约简方法都是精确的,即在约简后的数据集上学习到的模型与在完整数据集上学习到的模型是相同的。该项目在很大程度上依赖于优化理论,特别是灵敏度分析和凸几何。该项目的成果包括一个统一的方法来加速稀疏学习,并为开发高效和精确的数据约简方法提供一个系统的框架。对冗余数据识别的系统研究和深入探索有望加深对稀疏学习技术的理解,并显著增强其在大数据分析中的应用。
英文摘要
Research in big data involves analyzing growing data sets with huge numbers of samples, very high-dimensional feature vectors, and complex and diverse structures. The ever-growing volume and complexity of these data sets make many traditional techniques inadequate to extract knowledge from them. An emerging area, known as sparse learning, has achieved great success in learning from big data by identifying a small set of explanatory features and/or samples. Typical examples include selecting features that are most indicative of users? preferences for recommendation systems, identifying brain regions that are predictive of neurological disorders based on imaging data, and extracting semantic information from raw images for object recognition. However, training sparse learning models can be computationally prohibitive due to the sparsity-inducing regularization, which is non-smooth and can be highly complex when incorporating complex structures. This project aims at developing algorithms and tools to significantly accelerate the training process of sparse learning models for big data applications. The key idea is to efficiently identify redundant features and/or samples, which can be removed from the training phase without losing useful information of interests. Success in these unique techniques is expected to dramatically scaling up sparse learning for big data by orders of magnitude in terms of both time and space. The PIs plan to integrate the big data reduction tools developed in this project into their education and outreach activities, including development of new courses and integration of project components into existing courses. The PIs will make special efforts to recruit female and underrepresented students to this project.The major technical innovations of this project include the following components: (1) the PIs will develop efficient feature reduction methods for the generic scenario where the structures of both input and output can be represented by directed acyclic graphs; the proposed formulations include many existing approaches as special cases; (2) the PIs will develop efficient methods to reduce the numbers of features and samples simultaneously under a unified formulation, which can also incorporate various structures; (3) the PIs will develop efficient methods to discard irrelevant data subspaces to accelerate the process of uncovering low-rank structures commonly seen in big data. All the proposed data reduction methods are exact, i.e., the models learned on the reduced data sets are identical to the ones learned on the full data sets. This project heavily relies on optimization theory, especially on sensitivity analysis and convex geometry. The outcome of this project includes a unified approach to accelerate sparse learning and provide a systematic framework for developing efficient and exact data reduction methods. The systematic study and in-depth exploration of redundant data identification is expected to deepen the understanding of sparse learning techniques and dramatically enhance their applications in big data analytics.
期刊论文(30)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Explainable Prediction of Text Complexity: The Missing Preliminaries for Text Simplification
文本复杂性的可解释预测:文本简化缺失的预备知识
DOI:
10.18653/v1/2021.acl-long.88
发表时间:
2021
期刊:
Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing
影响因子:
--
作者:
[Garbacea, Cristina, Guo, Mengtian, Carton, Samuel, Mei, Qiaozhu]
通讯作者:
Mei, Qiaozhu
DOI:
10.1145/3425866
发表时间:
2021-03
期刊:
ACM Transactions on Internet Technology
影响因子:
5.3
作者:
[LiuXuanzhe;WangShangguang;MaYun;Zhangying;MeiQiaozhu;LiuYunxin;Huanggang]
通讯作者:
LiuXuanzhe;WangShangguang;MaYun;Zhangying;MeiQiaozhu;LiuYunxin;Huanggang
DOI:
10.18653/v1/d18-1386
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
作者:
[Samuel Carton;Qiaozhu Mei;P. Resnick]
通讯作者:
Samuel Carton;Qiaozhu Mei;P. Resnick
DOI:
--
发表时间:
2021-06
期刊:
Mathematical Programming
影响因子:
2.7
作者:
[Jiaqi Ma;Junwei Deng;Qiaozhu Mei]
通讯作者:
Jiaqi Ma;Junwei Deng;Qiaozhu Mei
Decoding the New World Language: Analyzing the Popularity, Roles, and Utility of Emojis
解码新世界语言:分析表情符号的流行度、作用和实用性
DOI:
10.1145/3308560.3316541
发表时间:
2019
期刊:
WWW '19 Companion: Companion Proceedings of the 2019 World Wide Web Conference
影响因子:
--
作者:
[Mei, Qiaozhu]
通讯作者:
Mei, Qiaozhu
共 23 条
NSF Student Travel Grant for 2017 Conference on Knowledge Discovery and Data Mining (KDD 2017)
-
批准号:1742808
-
项目类别:Standard Grant
-
资助金额:$2.5万
-
财政年份:2017
-
负责人:Qiaozhu Mei
-
依托单位:
CiC (RDDC): Wordsmith in the Cloud - Refining Language Models Using Web-Scale Language Networks
-
批准号:1048168
-
项目类别:Standard Grant
-
资助金额:$21.5万
-
财政年份:2011
-
负责人:Qiaozhu Mei
-
依托单位:
CAREER: Eyes of the Foreseer - Integrative and In Situ Information Retrieval and Mining in Online Communities
-
批准号:1054199
-
项目类别:Continuing Grant
-
资助金额:$44.95万
-
财政年份:2011
-
负责人:Qiaozhu Mei
-
依托单位:
SoCS: Assessing Information Credibility Without Authoritative Sources
-
批准号:0968489
-
项目类别:Standard Grant
-
资助金额:$75.0万
-
财政年份:2010
-
负责人:Qiaozhu Mei
-
依托单位:
海外基金