EAGER-DynamicData: A new paradigm for data analytics: L1-norm based Learning and Processing
EAGER-DynamicData: A new paradigm for data analytics: L1-norm based Learning and Processing
批准号:
1462341
负责人:
Michael Langberg
金额:
$18.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2018-08-31
中文摘要
该项目旨在就数据分析和数据特征提取的变革性/颠覆性思想开展基础性工作,新定义和计算的主成分向量最能代表给定数据集的主要特征,即使在存在错误/缺失/离群点数据的情况下也是如此。基于这一想法,将开发一个算法框架来支持几个有针对性的应用,如社交网络中的数据处理、图像处理和视频库、无线传感器网络数据融合、经济学、基因组学和蛋白质组学以及生物信息学。该项目工作的潜在影响是巨大的,可能远远超出这些应用程序,覆盖过去使用传统数据特征提取的任何科学和工程领域。从技术上讲,该研究旨在改写过去一个世纪关于L2范数(特征向量和奇异向量分解)数据分析的巨大回报的章节。正在开发最优的L1范数数据分析,这种数据分析天生就能抵抗数据污染,并与对“干净”数据的L2范数分析一样好。L1范数数据的主成分分析在以往的研究中数量有限,到目前为止在教育/教科书中还不存在。基于L1范数的主成分分析(PCA)和标准的L2范数主成分分析(PCA)之间的一些深刻差异,到目前为止阻碍了对基于L1范数的主成分分析的理论理解(从而在设计有效的算法解决方案方面)的进展。该项目寻求开发一种在L1规范下进行降维的新方法,以处理易出现离群点/受污染的大数据?(大量高维数据)通过将基本的L1范数主成分优化问题解释为等价的二元场最大化问题,并在这种情况下开辟了一系列潜在的新的分析和算法技术。项目目标包括:(1)关于最大L1范数投影数据特征精确计算的基本算法研究;(2)对L1范数测量数据降维的基本理解和执行;(3)样本空间降维。
英文摘要
This project aims to carry out fundamental work on the transformative/disruptive idea of data analytics and data-feature extraction by newly defined and calculated principal-component vectors that best represent the main features of a given data set, even in the presence of faulty/missing/outlier data. Based on this idea, an algorithmic framework will be developed to support several targeted applications, such as data processing in social networks, image processing and video libraries, wireless-sensor-network data fusion, economics, genomics and proteomics, and bioinformatics. The potential impact of the project work is immense and may extend well beyond these applications to cover any field of science and engineering where conventional data feature extraction has been used in the past.Technically speaking, the investigation aims at rewriting the enormously rewarding over the past century chapter on L2-norm (eigen-vector and singular-vector decomposition) data analysis. Optimal L1-norm data analytics are being developed that are inherently resistant to data contamination and as good as L2-norm analytics on "clean" data. L1-norm data principal component analysis has seen a limited amount of previous research and is non-existent so far in education/textbooks. Several profound differences between L1-norm based principle component analysis (PCA) and standard L2-norm PCA have, to date, blocked progress in the theoretical understanding (and thus in the design of efficient algorithmic solutions) of L1-based PCA. The project seeks the development of a novel approach toward dimensionality reduction under the L1 norm to deal with processing of outlier-prone/contaminated ?big data? (large amount of high-dimensional data) by interpreting the fundamental L1-norm principal-components optimization problem as an equivalent binary-field maximization problem, and in such opening a spectrum of potentially new analytical and algorithmic techniques. The project goals include: (i) Fundamental algorithmic research on the exact calculation of maximum-L1-norm-projection data features; (ii) fundamental understanding and execution of L1-norm-measured data dimensionality reduction; and (iii) sample-space reduction.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CIF: Small: Key Dissemination over Networks
-
批准号:2245204
-
项目类别:Standard Grant
-
资助金额:$55.8万
-
财政年份:2023
-
负责人:Michael Langberg
-
依托单位:
CIF: Small: Collaborative Research: Between Shannon and Hamming
-
批准号:1909451
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2019
-
负责人:Michael Langberg
-
依托单位:
CIF: Small: Collaborative Research:A Reductionist View of Network Information Theory
-
批准号:1526771
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2015
-
负责人:Michael Langberg
-
依托单位:
海外基金