课题基金 / 基金详情

Moment Invariant Data Aggregation for Signal Processing and Distribution Learning

Moment Invariant Data Aggregation for Signal Processing and Distribution Learning
用于信号处理和分布学习的矩不变数据聚合
批准号:
2309570
负责人:
Anna Little
金额:
$36.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-06-01 至 2026-05-31

项目摘要

项目成果

Anna Little的其他基金

相似基金

相关文献

中文摘要
翻译
在现代时代,来自不同来源的数据爆炸式增长大大增加了对可靠数据集成工具的需求。在许多应用程序中,可以访问大量数据,但是这些数据非常嘈杂,只有在聚合时才有意义。例如,在低温电子显微镜中,人们可以接触到大量的分子图像。尽管如此,每张图像都有很大的噪声,并且反映了随机的移动、方向和投影,这使得集成非常具有挑战性。又如微阵列表达数据,一般是小批量采集,批量效果显著。传统的统计方法是先“将数据标准化”再进行整合,即先用每批数据的标准差减去平均值,然后将标准化后的数据组合起来。然而,当每个批/亚总体的样本量很小时,这种方法将失败,因为批内均值和方差的估计可能不可靠。该项目将开发计算效率高且可扩展的方法,用于聚合来自多个来源的噪声数据。该方法将应用于刑事司法系统的数据,以深入了解基于种族和性别的差异。该项目将支持几个层次的指导以及研究生和本科生的研究,并将提供公开可用的代码。此外,该研究还将成为信号处理和统计学领域之间的桥梁。该项目旨在开发数学工具,用于在两种不同的环境中对数据的第一和第二时刻保持不变的数据聚合。第一个背景是多参考比对(MRA)的数据聚合,这是一个由低温电子显微镜等生物应用驱动的主题。在经典的MRA中,人们试图从许多有噪声的观测中恢复隐藏的信号,其中每个有噪声的观测都被随机翻译并被加性噪声破坏。研究人员将探索这个模型的泛化,其中每个嘈杂的观察也被随机的尺度变化所破坏。该研究将开发一种计算效率高的全信号恢复方法,该方法利用傅立叶和基于小波的特征来消除随机尺度变化的噪声数据的偏差。第二个上下文是用于分布学习的数据聚合。研究人员考虑了这样一种情况,即从不同的亚群体中收集数据以产生数据批次,但批次的样本量很小。在一个简单但令人信服的模型中,局部化因子仅影响子种群的第一阶矩和第二阶矩,因此每批子种群由一些移位和重新标度的普遍分布函数的独立、同分布的观测值组成。目标是通过聚合稀疏数据来恢复底层分布。这是一个高度相关的问题,因为它允许在传统方法失败的情况下对密度进行可靠的非参数估计,更广泛地说,可以在数据很少的情况下对亚种群进行更精确的比较。通过将该问题视为广义MRA问题,其中随机抽样的不确定性取代高斯噪声,该项目为其解决方案提供了一套创新的工具,灵感来自信号处理方法。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The explosion of data in the modern era from diverse sources has greatly increased the need for reliable data integration tools. In many applications, one has access to vast amounts of data, but the data is highly noisy and only meaningful when aggregated. For example, in cryo-electron microscopy, one has access to a large volume of molecular images. Still, each image is highly noisy and reflects a random shift, orientation, and projection, making integration highly challenging. Another example is microarray expression data, which is generally collected in small batches with significant batch effects. The traditional statistical approach is to “standardize the data” before integration, i.e., one subtracts the mean and scales by the standard deviation in each batch and then combines the standardized batches. However, when the sample size of each batch/subpopulation is small, this approach will fail since within-batch estimates of mean and variance can be unreliable. This project will develop computationally efficient and scalable methods for aggregating noisy data from multiple sources. The methodology will be applied to data from the criminal justice system to gain insights into race and gender-based discrepancies. The project will support several tiers of mentorship as well as graduate and undergraduate research and will contribute publicly available code. In addition, the research will also serve as a bridge between the fields of signal processing and statistics. The project aims to develop mathematical tools for data aggregation invariant to the first and second moments of the data in two distinct contexts. The first context is data aggregation for multi-reference alignment (MRA), a topic motivated by biological applications such as cryo-electron microscopy. In classic MRA one attempts to recover a hidden signal from many noisy observations, where each noisy observation has been randomly translated and corrupted by additive noise. The investigators will explore a generalization of this model where each noisy observation is also corrupted by a random scale change. The research will develop a computationally efficient method for full signal recovery that utilizes Fourier and wavelet-based features to unbias the noisy data for the random scale change. The second context is data aggregation for distribution learning. The investigators consider the scenario where data is collected from various sub-populations to produce data batches, but the sample sizes of the batches are small. Under a simple but compelling model, localization factors affect only the first and second moments of the sub-populations, so that each batch consists of independent, identically distributed observations from some shifted and rescaled universal distribution function. The goal is to recover the underlying distribution by aggregating the sparse data. This is a highly relevant problem, as it allows for reliable nonparametric estimates of the density in settings where traditional approaches fail and, more broadly, leads to more precise comparisons across sub-populations in settings where little data is available. By viewing this problem as a generalized MRA problem where uncertainty due to random sampling replaces Gaussian noise, this project contributes an innovative set of tools for its solution inspired by methods in signal processing.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Data-driven Path Metrics for Machine Learning
  • 批准号:
    2131292
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2021
  • 负责人:
    Anna Little
  • 依托单位:
Collaborative Research: Data-driven Path Metrics for Machine Learning
  • 批准号:
    1912906
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2019
  • 负责人:
    Anna Little
  • 依托单位:
海外基金