BIGDATA: F: Big Data Analysis via Non-Standard Property Testing
BIGDATA: F: Big Data Analysis via Non-Standard Property Testing
批准号:
1838154
负责人:
Rocco Servedio
金额:
$91.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-01-01 至 2023-12-31
中文摘要
在现代,各个领域都在不断产生海量数据:包括正在进行的大规模科学实验,无处不在的智能手机和传感器,社交媒体上内容的持续生产和演变,以及许多其他领域。如何有效地处理和分析这些海量数据?计算机科学的一个分支“属性测试”寻求开发超快算法来分析海量数据集,以快速确定数据是否具有某些感兴趣的属性。然而,在性能测试中主要考虑的标准理论模型并不能很好地适用于许多真实世界的数据分析场景;这些标准模型优先考虑数学优雅,但它们所做的假设与实际数据分析算法的能力或许多实际数据集的性质不太一致。(作为一个例子,这些模型通常假设数据分析算法可以合成任意数据点并对其进行查询以接收关于应该如何标记这些数据点的准确信息,但是在许多真实世界的设置中,这种查询是不可能的,在这些环境中,数据点是按原样出现的,并且不能被合成以满足数据分析师的规范。作为另一个例子,这些模型通常只能处理被假设服从某些高度结构化的概率分布的数据,但现实世界的数据是杂乱无章的,很少具有如此高度的结构。)这个项目的高级目标是开发和分析非标准的性能测试模型,明确的目标是开发与现实世界数据分析问题的现实和约束相一致的算法。一个重要的相关目标是,通过开展外联活动和培训研究生,包括历来任职人数不足的群体的成员,促进人力资源开发,使其掌握对该项目至关重要的分析和算法技术。为了实现更广泛的影响,计划的活动还包括新课程、调查文章,以及继续针对中小学生的推广活动。更详细地说,该项目将专注于大数据属性测试算法的三个不同方面,所有这些都是出于对真实世界数据分析的考虑:(1)该项目的第一个重点将是开发灵活的算法,用于测试海量高维数据集是否已被标记为“军政府”-这是一种标记规则,它只依赖于巨大的可能特征集中非常小但未知的数据特征集。在之前工作的基础上,调查人员将致力于开发军政府测试算法,这种算法可以处理任意数据分布和噪声数据,即使只给出对正在分析的数据集的有限形式的访问,也可以成功。(2)该项目的第二个重点将是将思想和技术从理论机器学习算法转移到海量数据集的性能测试领域。研究人员之前的工作为如何修改(某些相对低效的)机器学习算法以产生用于数据分析的更有效的属性测试算法提供了概念验证,但这种转移仅在上文提到的相对受限的属性测试标准模型中进行,这些模型假设高度结构化的数据分布。在这个项目中,研究人员将致力于扩展这些早期的结果,以便机器学习技术将产生更灵活的、具有更大现实适用性的属性测试模型的算法。(3)最后,本项目的第三个重点是开发不需要对合成数据点进行查询而只使用随机样本的属性测试算法,该算法可以应用于高维连续数据集。这种类型的数据通常出现在传感器或不同类型的测量生成数据的环境中,但大多数属性测试算法是为离散二进制值数据而不是连续数据设计的。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In the modern era truly enormous amounts of data are constantly being generated across a wide range of domains: these include ongoing large-scale scientific experiments, ubiquitous smartphones and sensors, the continuous production and evolution of content on social media, and many others. How can this flood of data be efficiently processed and analyzed? A branch of computer science called "property testing" seeks to develop ultra-fast algorithms for analyzing massive data sets to quickly determine whether or not the data has some property of interest. However, the standard theoretical models that have mostly been considered in property testing are not well suited to many real-world data analysis scenarios; these standard models prioritize mathematical elegance, but the resulting assumptions they make do not align well with the abilities of actual data analysis algorithms or with the nature of many actual data sets. (As one example, these models typically assume that a data analysis algorithm can synthesize arbitrary data points and query them to receive accurate information about how such data points should be labeled, but such queries are impossible in many real-world settings where data points "come as they are" and cannot be synthesized to meet the specifications of a data analyst. As another example, these models typically can only deal with data which is assumed to follow certain highly structured probability distributions, but real-world data is messy and rarely possesses such a high degree of structure.) The high-level goal of this project is to develop and analyze non-standard models of property testing, with the explicit goal of developing algorithms which align with the realities and constraints of real-world data analysis problems. An important related goal is to foster human resource development by performing outreach and training graduate students, including members of historically under-represented groups, in the analytic and algorithmic techniques that are central to this project. Planned activities to achieve broader impacts also include new courses, survey articles, and the continuation of outreach activities aimed at students at the elementary and middle school levels.In more detail, the project will focus on three different aspects of property testing algorithms for big data, all of which are motivated by considerations arising from real-world data analysis:(1) The first focus of the project will be on developing flexible algorithms for testing whether a massive high-dimensional data set has been labeled according to a "junta" --- this is a labeling rule which depends only on a very small but unknown set of data features out of a huge set of possible features. Building on their previous work, the investigators will work to develop junta testing algorithms which can handle arbitrary data distributions and noisy data, and can succeed even given only a limited form of access to the data set being analyzed. (2) The second focus of the project will be on transferring ideas and techniques from theoretical machine learning algorithms to the domain of property testing of massive data sets. Previous work of the investigators gave a proof-of-concept for how (certain relatively inefficient) machine learning algorithms can be modified to yield far more efficient property testing algorithms for data analysis, but this transfer went through only in the relatively constrained standard models of property testing, alluded to above, which assume highly structured data distributions. In this project the investigators will work to extend these earlier results so that the machine learning techniques will yield algorithms for more flexible property testing models that are of greater real-world applicability.(3) Finally, the third focus of the project is to develop property testing algorithms which do not need to make queries on synthetic data points but instead use only random samples, and which can be applied to high-dimensional continuous data sets. Data of this type arises commonly in settings where sensors or measurements of different sorts are generating the data, but most property testing algorithms are designed for discrete binary-valued data rather than continuous data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(18)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Near-Optimal Average-Case Approximate Trace Reconstruction from Few Traces
从少量迹线重建近乎最优的平均情况近似迹线
DOI:
--
发表时间:
2022
期刊:
Proceedings of the annual ACMSIAM symposium on discrete algorithms
影响因子:
--
作者:
[Chen, Xi, De, Anindya, Lee, Chin Ho, Servedio, Rocco A., Sinha, Sandip]
通讯作者:
Sinha, Sandip
Random Restrictions of High-Dimensional Distributions and Uniformity Testing with Subcube Conditioning
高维分布的随机限制和子立方条件的均匀性测试
DOI:
10.5555/3458064.3458085
发表时间:
2021
期刊:
Proceedings of the 32th Annual ACM-SIAM Symposium on Discrete Algorithms
影响因子:
--
作者:
[Canonne, Clement, Chen, Xi, Kamath, Gautam, Levi, Amit, Waingarten, Erik]
通讯作者:
Waingarten, Erik
DOI:
10.1145/3519935.3519979
发表时间:
2021-11
期刊:
Proceedings of the 54th Annual ACM SIGACT Symposium on Theory of Computing
影响因子:
--
作者:
[Xi Chen;Rajesh Jayaram;Amit Levi;Erik Waingarten]
通讯作者:
Xi Chen;Rajesh Jayaram;Amit Levi;Erik Waingarten
Approximating Sumset Size
近似总集大小
DOI:
10.1137/1.9781611977073.94
发表时间:
2022
期刊:
ACM-SIAM Symposium on Discrete Algorithms
影响因子:
--
作者:
[De, Anindya, Nadimpali, Shivam, Servedio, Rocco A.]
通讯作者:
Servedio, Rocco A.
DOI:
--
发表时间:
2023
期刊:
Proceedings of the 2023 {ACM-SIAM} Symposium on Discrete Algorithms
影响因子:
--
作者:
[De, Anindya, Nadimpali, Shivam, Servedio, Rocco A.]
通讯作者:
Servedio, Rocco A.
共 16 条
Collaborative Research: AF: Medium: Continuous Concrete Complexity
-
批准号:2211238
-
项目类别:Continuing Grant
-
资助金额:$60.0万
-
财政年份:2022
-
负责人:Rocco Servedio
-
依托单位:
AF: Medium: The Trace Reconstruction Problem
-
批准号:2106429
-
项目类别:Continuing Grant
-
资助金额:$120.0万
-
财政年份:2021
-
负责人:Rocco Servedio
-
依托单位:
NSF QCIS-FF: Columbia University Computer Science Department Proposal
-
批准号:1926524
-
项目类别:Continuing Grant
-
资助金额:$75.0万
-
财政年份:2020
-
负责人:Rocco Servedio
-
依托单位:
Student Travel Grant for 2019 Conference on Computational Complexity (CCC)
-
批准号:1919026
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2019
-
负责人:Rocco Servedio
-
依托单位:
AF: Small: Collaborative Research: Boolean Function Analysis Meets Stochastic Design
-
批准号:1814873
-
项目类别:Standard Grant
-
资助金额:$16.63万
-
财政年份:2018
-
负责人:Rocco Servedio
-
依托单位:
Student Travel Support for CCC 2018
-
批准号:1822097
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:2018
-
负责人:Rocco Servedio
-
依托单位:
AF: Student Travel to CCC 2017
-
批准号:1724073
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:2017
-
负责人:Rocco Servedio
-
依托单位:
AF: Medium: Collaborative Research: Circuit Lower Bounds via Projections
-
批准号:1563155
-
项目类别:Continuing Grant
-
资助金额:$84.15万
-
财政年份:2016
-
负责人:Rocco Servedio
-
依托单位:
AF: Small: Linear and Polynomial Threshold Functions: Structural Analysis and Algorithmic Applications
-
批准号:1420349
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2014
-
负责人:Rocco Servedio
-
依托单位:
AF: Small: Learning and Testing Classes of Distributions
-
批准号:1319788
-
项目类别:Standard Grant
-
资助金额:$47.19万
-
财政年份:2013
-
负责人:Rocco Servedio
-
依托单位:
Student Travel to STOC 2013
-
批准号:1319775
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:2013
-
负责人:Rocco Servedio
-
依托单位:
AF: Small: The Boundary of Learnability for Monotone Boolean Functions
-
批准号:1115703
-
项目类别:Standard Grant
-
资助金额:$35.0万
-
财政年份:2011
-
负责人:Rocco Servedio
-
依托单位:
AF: Small: Collaborative Research: The Polynomial Method for Learning
-
批准号:0915929
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2009
-
负责人:Rocco Servedio
-
依托单位:
CT-ISG: Cross-Leveraging Cryptography with Learning Theory
-
批准号:0716245
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Rocco Servedio
-
依托单位:
QnTM: Quantum Computational Learning
-
批准号:0523664
-
项目类别:Continuing Grant
-
资助金额:$28.0万
-
财政年份:2005
-
负责人:Rocco Servedio
-
依托单位:
CAREER: Efficient Learning Algorithms for Rich Function Classes
-
批准号:0347282
-
项目类别:Continuing Grant
-
资助金额:$40.0万
-
财政年份:2004
-
负责人:Rocco Servedio
-
依托单位:
Efficient Algorithms in Computational Learning Theory
-
批准号:0102075
-
项目类别:Fellowship Award
-
资助金额:$9.0万
-
财政年份:2001
-
负责人:Rocco Servedio
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
ARF鸟苷酸交换因子BIG1介导ACSL4依赖性铁死亡在非酒精性脂肪性肝炎中的作用及机制研究
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:游艳
-
依托单位:
基于Big Code深度背景增强的Android应用代码反混淆研究
-
批准号:61972290
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2019
-
负责人:刘进
-
依托单位:
BIG1介导STING囊泡转运在抗肺癌免疫反应中的作用及分子机制
-
批准号:81903639
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:张素林
-
依托单位:
水稻Big Grain3 通过调控细胞分裂素转运调节籽粒大小
-
批准号:2019JJ50243
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2019
-
负责人:肖云华
-
依托单位:
ARF鸟苷酸交换因子BIG1调控巨噬细胞重编程在脓毒症免疫抑制形成中的作用及机制研究
-
批准号:81971488
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2019
-
负责人:沈晓燕
-
依托单位:
控制豆科作物器官大小关键基因BIG SEEDS1的功能与应用研究
-
批准号:31771345
-
项目类别:面上项目
-
资助金额:65.0万元
-
批准年份:2017
-
负责人:葛良法
-
依托单位:
生长素转运调控基因BIG介导高浓度CO2下气孔关闭的分子机制
-
批准号:31171356
-
项目类别:面上项目
-
资助金额:65.0万元
-
批准年份:2011
-
负责人:梁允宽
-
依托单位:
ARF鸟苷酸交换因子BIG1定向调控ABCA1功能的分子机制
-
批准号:81173056
-
项目类别:面上项目
-
资助金额:69.0万元
-
批准年份:2011
-
负责人:沈晓燕
-
依托单位:
BIG2介导的GABAA型受体转运模式及信号调控机制
-
批准号:31070924
-
项目类别:面上项目
-
资助金额:35.0万元
-
批准年份:2010
-
负责人:沈晓燕
-
依托单位: