BIGDATA: F: Big Data Analysis via Non-Standard Property Testing
BIGDATA: F: Big Data Analysis via Non-Standard Property Testing
批准号:
1838154
负责人:
Rocco Servedio
金额:
$91.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-01-01 至 2023-12-31
中文摘要
在当今时代,在广泛的领域中不断产生真正巨大的数据量:这些包括正在进行的大规模科学实验,无处不在的智能手机和传感器,社交媒体上内容的持续生产和演变等等。 如何有效地处理和分析这些海量数据? 计算机科学的一个分支称为“属性测试”,旨在开发用于分析大量数据集的超快速算法,以快速确定数据是否具有某些感兴趣的属性。 然而,大多数在属性测试中考虑的标准理论模型并不适合许多现实世界的数据分析场景;这些标准模型优先考虑数学优雅,但它们所做的假设与实际数据分析算法的能力或许多实际数据集的性质并不一致。 (As在一个示例中,这些模型通常假设数据分析算法可以合成任意数据点,并查询它们以接收关于这些数据点应该如何被标记的准确信息,但是这样的查询在数据点“原样出现”的许多现实世界设置中是不可能的,并且不能被合成以满足数据分析师的规范。 作为另一个例子,这些模型通常只能处理假设遵循某些高度结构化概率分布的数据,但现实世界的数据是混乱的,很少具有如此高的结构度。 该项目的高级目标是开发和分析属性测试的非标准模型,其明确目标是开发与现实世界数据分析问题的现实和约束相一致的算法。 一个重要的相关目标是促进人力资源开发,方法是对研究生进行外联和培训,包括对历史上代表性不足的群体的成员进行分析和算法技术方面的培训,这对该项目至关重要。 为了实现更广泛的影响,计划开展的活动还包括新课程,调查文章,以及针对小学和中学学生的持续外展活动。更详细地说,该项目将专注于大数据属性测试算法的三个不同方面,所有这些都是出于对真实世界数据分析的考虑:(1)该项目的第一个重点将是开发灵活的算法,用于测试是否已根据“军政府”标记了大量高维数据集-这是一种标记规则,其仅依赖于可能特征的巨大集合中的非常小但未知的数据特征集合。 在他们以前工作的基础上,研究人员将致力于开发军政府测试算法,该算法可以处理任意数据分布和噪声数据,即使只有有限的数据集访问形式也可以成功分析。(2)该项目的第二个重点是将理论机器学习算法的思想和技术转移到海量数据集的属性测试领域。 研究人员以前的工作给出了一个概念验证,说明如何修改(某些相对低效的)机器学习算法,以产生更有效的数据分析属性测试算法,但这种转移只在相对受限的属性测试标准模型中进行,上面提到,假设高度结构化的数据分布。 在这个项目中,研究人员将努力扩展这些早期的结果,以便机器学习技术将产生更灵活的属性测试模型的算法,这些模型具有更大的现实适用性。(3)最后,该项目的第三个重点是开发属性测试算法,这些算法不需要对合成数据点进行查询,而是只使用随机样本,并且可以应用于高维连续数据集。 这种类型的数据通常出现在传感器或不同种类的测量产生数据的设置中,但大多数属性测试算法是为离散二进制值数据而不是连续数据设计的。该奖项反映了NSF的法定使命,并被认为值得通过使用基金会的知识价值和更广泛的影响审查标准进行评估来支持。
英文摘要
In the modern era truly enormous amounts of data are constantly being generated across a wide range of domains: these include ongoing large-scale scientific experiments, ubiquitous smartphones and sensors, the continuous production and evolution of content on social media, and many others. How can this flood of data be efficiently processed and analyzed? A branch of computer science called "property testing" seeks to develop ultra-fast algorithms for analyzing massive data sets to quickly determine whether or not the data has some property of interest. However, the standard theoretical models that have mostly been considered in property testing are not well suited to many real-world data analysis scenarios; these standard models prioritize mathematical elegance, but the resulting assumptions they make do not align well with the abilities of actual data analysis algorithms or with the nature of many actual data sets. (As one example, these models typically assume that a data analysis algorithm can synthesize arbitrary data points and query them to receive accurate information about how such data points should be labeled, but such queries are impossible in many real-world settings where data points "come as they are" and cannot be synthesized to meet the specifications of a data analyst. As another example, these models typically can only deal with data which is assumed to follow certain highly structured probability distributions, but real-world data is messy and rarely possesses such a high degree of structure.) The high-level goal of this project is to develop and analyze non-standard models of property testing, with the explicit goal of developing algorithms which align with the realities and constraints of real-world data analysis problems. An important related goal is to foster human resource development by performing outreach and training graduate students, including members of historically under-represented groups, in the analytic and algorithmic techniques that are central to this project. Planned activities to achieve broader impacts also include new courses, survey articles, and the continuation of outreach activities aimed at students at the elementary and middle school levels.In more detail, the project will focus on three different aspects of property testing algorithms for big data, all of which are motivated by considerations arising from real-world data analysis:(1) The first focus of the project will be on developing flexible algorithms for testing whether a massive high-dimensional data set has been labeled according to a "junta" --- this is a labeling rule which depends only on a very small but unknown set of data features out of a huge set of possible features. Building on their previous work, the investigators will work to develop junta testing algorithms which can handle arbitrary data distributions and noisy data, and can succeed even given only a limited form of access to the data set being analyzed. (2) The second focus of the project will be on transferring ideas and techniques from theoretical machine learning algorithms to the domain of property testing of massive data sets. Previous work of the investigators gave a proof-of-concept for how (certain relatively inefficient) machine learning algorithms can be modified to yield far more efficient property testing algorithms for data analysis, but this transfer went through only in the relatively constrained standard models of property testing, alluded to above, which assume highly structured data distributions. In this project the investigators will work to extend these earlier results so that the machine learning techniques will yield algorithms for more flexible property testing models that are of greater real-world applicability.(3) Finally, the third focus of the project is to develop property testing algorithms which do not need to make queries on synthetic data points but instead use only random samples, and which can be applied to high-dimensional continuous data sets. Data of this type arises commonly in settings where sensors or measurements of different sorts are generating the data, but most property testing algorithms are designed for discrete binary-valued data rather than continuous data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(18)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Near-Optimal Average-Case Approximate Trace Reconstruction from Few Traces
从少量迹线重建近乎最优的平均情况近似迹线
DOI:
--
发表时间:
2022
期刊:
Proceedings of the annual ACMSIAM symposium on discrete algorithms
影响因子:
--
作者:
[Chen, Xi, De, Anindya, Lee, Chin Ho, Servedio, Rocco A., Sinha, Sandip]
通讯作者:
Sinha, Sandip
Random Restrictions of High-Dimensional Distributions and Uniformity Testing with Subcube Conditioning
高维分布的随机限制和子立方条件的均匀性测试
DOI:
10.5555/3458064.3458085
发表时间:
2021
期刊:
Proceedings of the 32th Annual ACM-SIAM Symposium on Discrete Algorithms
影响因子:
--
作者:
[Canonne, Clement, Chen, Xi, Kamath, Gautam, Levi, Amit, Waingarten, Erik]
通讯作者:
Waingarten, Erik
DOI:
10.1145/3519935.3519979
发表时间:
2021-11
期刊:
Proceedings of the 54th Annual ACM SIGACT Symposium on Theory of Computing
影响因子:
--
作者:
[Xi Chen;Rajesh Jayaram;Amit Levi;Erik Waingarten]
通讯作者:
Xi Chen;Rajesh Jayaram;Amit Levi;Erik Waingarten
Approximating Sumset Size
近似总集大小
DOI:
10.1137/1.9781611977073.94
发表时间:
2022
期刊:
ACM-SIAM Symposium on Discrete Algorithms
影响因子:
--
作者:
[De, Anindya, Nadimpali, Shivam, Servedio, Rocco A.]
通讯作者:
Servedio, Rocco A.
DOI:
--
发表时间:
2023
期刊:
Proceedings of the 2023 {ACM-SIAM} Symposium on Discrete Algorithms
影响因子:
--
作者:
[De, Anindya, Nadimpali, Shivam, Servedio, Rocco A.]
通讯作者:
Servedio, Rocco A.
共 16 条
Collaborative Research: AF: Medium: Continuous Concrete Complexity
-
批准号:2211238
-
项目类别:Continuing Grant
-
资助金额:$60.0万
-
财政年份:2022
-
负责人:Rocco Servedio
-
依托单位:
AF: Medium: The Trace Reconstruction Problem
-
批准号:2106429
-
项目类别:Continuing Grant
-
资助金额:$120.0万
-
财政年份:2021
-
负责人:Rocco Servedio
-
依托单位:
NSF QCIS-FF: Columbia University Computer Science Department Proposal
-
批准号:1926524
-
项目类别:Continuing Grant
-
资助金额:$75.0万
-
财政年份:2020
-
负责人:Rocco Servedio
-
依托单位:
Student Travel Grant for 2019 Conference on Computational Complexity (CCC)
-
批准号:1919026
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2019
-
负责人:Rocco Servedio
-
依托单位:
AF: Small: Collaborative Research: Boolean Function Analysis Meets Stochastic Design
-
批准号:1814873
-
项目类别:Standard Grant
-
资助金额:$16.63万
-
财政年份:2018
-
负责人:Rocco Servedio
-
依托单位:
Student Travel Support for CCC 2018
-
批准号:1822097
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:2018
-
负责人:Rocco Servedio
-
依托单位:
AF: Student Travel to CCC 2017
-
批准号:1724073
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:2017
-
负责人:Rocco Servedio
-
依托单位:
AF: Medium: Collaborative Research: Circuit Lower Bounds via Projections
-
批准号:1563155
-
项目类别:Continuing Grant
-
资助金额:$84.15万
-
财政年份:2016
-
负责人:Rocco Servedio
-
依托单位:
AF: Small: Linear and Polynomial Threshold Functions: Structural Analysis and Algorithmic Applications
-
批准号:1420349
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2014
-
负责人:Rocco Servedio
-
依托单位:
AF: Small: Learning and Testing Classes of Distributions
-
批准号:1319788
-
项目类别:Standard Grant
-
资助金额:$47.19万
-
财政年份:2013
-
负责人:Rocco Servedio
-
依托单位:
Student Travel to STOC 2013
-
批准号:1319775
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:2013
-
负责人:Rocco Servedio
-
依托单位:
AF: Small: The Boundary of Learnability for Monotone Boolean Functions
-
批准号:1115703
-
项目类别:Standard Grant
-
资助金额:$35.0万
-
财政年份:2011
-
负责人:Rocco Servedio
-
依托单位:
AF: Small: Collaborative Research: The Polynomial Method for Learning
-
批准号:0915929
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2009
-
负责人:Rocco Servedio
-
依托单位:
CT-ISG: Cross-Leveraging Cryptography with Learning Theory
-
批准号:0716245
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Rocco Servedio
-
依托单位:
QnTM: Quantum Computational Learning
-
批准号:0523664
-
项目类别:Continuing Grant
-
资助金额:$28.0万
-
财政年份:2005
-
负责人:Rocco Servedio
-
依托单位:
CAREER: Efficient Learning Algorithms for Rich Function Classes
-
批准号:0347282
-
项目类别:Continuing Grant
-
资助金额:$40.0万
-
财政年份:2004
-
负责人:Rocco Servedio
-
依托单位:
Efficient Algorithms in Computational Learning Theory
-
批准号:0102075
-
项目类别:Fellowship Award
-
资助金额:$9.0万
-
财政年份:2001
-
负责人:Rocco Servedio
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
ARF鸟苷酸交换因子BIG1介导ACSL4依赖性铁死亡在非酒精性脂肪性肝炎中的作用及机制研究
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:游艳
-
依托单位:
基于Big Code深度背景增强的Android应用代码反混淆研究
-
批准号:61972290
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2019
-
负责人:刘进
-
依托单位:
BIG1介导STING囊泡转运在抗肺癌免疫反应中的作用及分子机制
-
批准号:81903639
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:张素林
-
依托单位:
水稻Big Grain3 通过调控细胞分裂素转运调节籽粒大小
-
批准号:2019JJ50243
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2019
-
负责人:肖云华
-
依托单位:
ARF鸟苷酸交换因子BIG1调控巨噬细胞重编程在脓毒症免疫抑制形成中的作用及机制研究
-
批准号:81971488
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2019
-
负责人:沈晓燕
-
依托单位:
控制豆科作物器官大小关键基因BIG SEEDS1的功能与应用研究
-
批准号:31771345
-
项目类别:面上项目
-
资助金额:65.0万元
-
批准年份:2017
-
负责人:葛良法
-
依托单位:
生长素转运调控基因BIG介导高浓度CO2下气孔关闭的分子机制
-
批准号:31171356
-
项目类别:面上项目
-
资助金额:65.0万元
-
批准年份:2011
-
负责人:梁允宽
-
依托单位:
ARF鸟苷酸交换因子BIG1定向调控ABCA1功能的分子机制
-
批准号:81173056
-
项目类别:面上项目
-
资助金额:69.0万元
-
批准年份:2011
-
负责人:沈晓燕
-
依托单位:
BIG2介导的GABAA型受体转运模式及信号调控机制
-
批准号:31070924
-
项目类别:面上项目
-
资助金额:35.0万元
-
批准年份:2010
-
负责人:沈晓燕
-
依托单位: