课题基金 / 基金详情

BIGDATA: F: Big Data Analysis via Non-Standard Property Testing

BIGDATA: F: Big Data Analysis via Non-Standard Property Testing
BIGDATA:F:通过非标准属性测试进行大数据分析
批准号:
1838154
负责人:
Rocco Servedio
金额:
$91.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-01-01 至 2023-12-31

项目摘要

项目成果

Rocco Servedio的其他基金

相似基金

相关文献

中文摘要
翻译
在现代,真正巨大的数据量在广泛的领域不断产生:这些领域包括正在进行的大规模科学实验,无处不在的智能手机和传感器,社交媒体上内容的持续生产和演变,以及许多其他领域。如何有效地处理和分析这些大量的数据?计算机科学的一个分支名为“属性测试”,旨在开发超快速算法,用于分析大量数据集,以快速确定数据是否具有某些感兴趣的属性。然而,在性能测试中主要考虑的标准理论模型并不适合许多现实世界的数据分析场景;这些标准模型优先考虑数学上的优雅,但是它们做出的最终假设与实际数据分析算法的能力或许多实际数据集的性质不太一致。(例如,这些模型通常假设数据分析算法可以合成任意数据点并对其进行查询,以接收有关如何标记这些数据点的准确信息,但在许多真实世界的设置中,这样的查询是不可能的,因为数据点“如其所是”,无法合成以满足数据分析师的规范。另一个例子是,这些模型通常只能处理假设遵循某种高度结构化概率分布的数据,但现实世界的数据是混乱的,很少具有如此高度的结构。)该项目的高级目标是开发和分析非标准的属性测试模型,其明确目标是开发与现实世界数据分析问题的现实和约束相一致的算法。一个重要的相关目标是通过开展拓展和培训研究生,包括历史上代表性不足的群体的成员,在分析和算法技术方面促进人力资源开发,这是该项目的核心。为取得更广泛影响而计划的活动还包括新课程、调查文章和继续开展针对中小学生的外联活动。更详细地说,该项目将侧重于大数据属性测试算法的三个不同方面,(1)该项目的第一个重点将是开发灵活的算法,用于测试大规模高维数据集是否已根据“军政府”进行了标记——这是一种标记规则,仅依赖于大量可能特征中的一组非常小但未知的数据特征。在他们之前工作的基础上,研究人员将致力于开发可以处理任意数据分布和噪声数据的测试算法,并且即使只提供对正在分析的数据集的有限形式的访问,也可以成功。(2)该项目的第二个重点将是将理论机器学习算法的思想和技术转移到大规模数据集的属性测试领域。研究人员之前的工作给出了一个概念证明,即如何修改(某些相对低效的)机器学习算法,以产生更有效的数据分析属性测试算法,但这种转移只发生在相对受限的属性测试标准模型中,如上所述,这些模型假设高度结构化的数据分布。在这个项目中,研究人员将努力扩展这些早期的结果,以便机器学习技术将产生更灵活的性能测试模型算法,这些模型在现实世界中具有更大的适用性。(3)最后,项目的第三个重点是开发不需要对合成数据点进行查询而只使用随机样本的属性测试算法,该算法可以应用于高维连续数据集。这种类型的数据通常出现在不同类型的传感器或测量产生数据的设置中,但大多数属性测试算法是为离散二值数据而不是连续数据设计的。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In the modern era truly enormous amounts of data are constantly being generated across a wide range of domains: these include ongoing large-scale scientific experiments, ubiquitous smartphones and sensors, the continuous production and evolution of content on social media, and many others. How can this flood of data be efficiently processed and analyzed? A branch of computer science called "property testing" seeks to develop ultra-fast algorithms for analyzing massive data sets to quickly determine whether or not the data has some property of interest. However, the standard theoretical models that have mostly been considered in property testing are not well suited to many real-world data analysis scenarios; these standard models prioritize mathematical elegance, but the resulting assumptions they make do not align well with the abilities of actual data analysis algorithms or with the nature of many actual data sets. (As one example, these models typically assume that a data analysis algorithm can synthesize arbitrary data points and query them to receive accurate information about how such data points should be labeled, but such queries are impossible in many real-world settings where data points "come as they are" and cannot be synthesized to meet the specifications of a data analyst. As another example, these models typically can only deal with data which is assumed to follow certain highly structured probability distributions, but real-world data is messy and rarely possesses such a high degree of structure.) The high-level goal of this project is to develop and analyze non-standard models of property testing, with the explicit goal of developing algorithms which align with the realities and constraints of real-world data analysis problems. An important related goal is to foster human resource development by performing outreach and training graduate students, including members of historically under-represented groups, in the analytic and algorithmic techniques that are central to this project. Planned activities to achieve broader impacts also include new courses, survey articles, and the continuation of outreach activities aimed at students at the elementary and middle school levels.In more detail, the project will focus on three different aspects of property testing algorithms for big data, all of which are motivated by considerations arising from real-world data analysis:(1) The first focus of the project will be on developing flexible algorithms for testing whether a massive high-dimensional data set has been labeled according to a "junta" --- this is a labeling rule which depends only on a very small but unknown set of data features out of a huge set of possible features. Building on their previous work, the investigators will work to develop junta testing algorithms which can handle arbitrary data distributions and noisy data, and can succeed even given only a limited form of access to the data set being analyzed. (2) The second focus of the project will be on transferring ideas and techniques from theoretical machine learning algorithms to the domain of property testing of massive data sets. Previous work of the investigators gave a proof-of-concept for how (certain relatively inefficient) machine learning algorithms can be modified to yield far more efficient property testing algorithms for data analysis, but this transfer went through only in the relatively constrained standard models of property testing, alluded to above, which assume highly structured data distributions. In this project the investigators will work to extend these earlier results so that the machine learning techniques will yield algorithms for more flexible property testing models that are of greater real-world applicability.(3) Finally, the third focus of the project is to develop property testing algorithms which do not need to make queries on synthetic data points but instead use only random samples, and which can be applied to high-dimensional continuous data sets. Data of this type arises commonly in settings where sensors or measurements of different sorts are generating the data, but most property testing algorithms are designed for discrete binary-valued data rather than continuous data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(18)
专著(0)
科研奖励(0)
会议论文
Near-Optimal Average-Case Approximate Trace Reconstruction from Few Traces
从少量迹线重建近乎最优的平均情况近似迹线
DOI: --
发表时间: 2022
期刊: Proceedings of the annual ACMSIAM symposium on discrete algorithms
影响因子: --
作者: [Chen, Xi, De, Anindya, Lee, Chin Ho, Servedio, Rocco A., Sinha, Sandip]
通讯作者: Sinha, Sandip
DOI: 10.5555/3458064.3458085
发表时间: 2021
期刊: Proceedings of the 32th Annual ACM-SIAM Symposium on Discrete Algorithms
影响因子: --
作者: [Canonne, Clement, Chen, Xi, Kamath, Gautam, Levi, Amit, Waingarten, Erik]
通讯作者: Waingarten, Erik
DOI: 10.1145/3519935.3519979
发表时间: 2021-11
期刊: Proceedings of the 54th Annual ACM SIGACT Symposium on Theory of Computing
影响因子: --
作者: [Xi Chen;Rajesh Jayaram;Amit Levi;Erik Waingarten]
通讯作者: Xi Chen;Rajesh Jayaram;Amit Levi;Erik Waingarten
Approximating Sumset Size
近似总集大小
DOI: 10.1137/1.9781611977073.94
发表时间: 2022
期刊: ACM-SIAM Symposium on Discrete Algorithms
影响因子: --
作者: [De, Anindya, Nadimpali, Shivam, Servedio, Rocco A.]
通讯作者: Servedio, Rocco A.
共 16 条
    Collaborative Research: AF: Medium: Continuous Concrete Complexity
    • 批准号:
      2211238
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $60.0万
    • 财政年份:
      2022
    • 负责人:
      Rocco Servedio
    • 依托单位:
    AF: Medium: The Trace Reconstruction Problem
    • 批准号:
      2106429
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $120.0万
    • 财政年份:
      2021
    • 负责人:
      Rocco Servedio
    • 依托单位:
    NSF QCIS-FF: Columbia University Computer Science Department Proposal
    • 批准号:
      1926524
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $75.0万
    • 财政年份:
      2020
    • 负责人:
      Rocco Servedio
    • 依托单位:
    Student Travel Grant for 2019 Conference on Computational Complexity (CCC)
    • 批准号:
      1919026
    • 项目类别:
      Standard Grant
    • 资助金额:
      $1.0万
    • 财政年份:
      2019
    • 负责人:
      Rocco Servedio
    • 依托单位:
    国内基金
    海外基金
    Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
    ARF鸟苷酸交换因子BIG1介导ACSL4依赖性铁死亡在非酒精性脂肪性肝炎中的作用及机制研究
    • 批准号:
      --
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      30万元
    • 批准年份:
      2022
    • 负责人:
      游艳
    • 依托单位:
    基于Big Code深度背景增强的Android应用代码反混淆研究
    • 批准号:
      61972290
    • 项目类别:
      面上项目
    • 资助金额:
      60.0万元
    • 批准年份:
      2019
    • 负责人:
      刘进
    • 依托单位:
    BIG1介导STING囊泡转运在抗肺癌免疫反应中的作用及分子机制
    • 批准号:
      81903639
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      21.0万元
    • 批准年份:
      2019
    • 负责人:
      张素林
    • 依托单位: