课题基金 / 基金详情

Moment Invariant Data Aggregation for Signal Processing and Distribution Learning

Moment Invariant Data Aggregation for Signal Processing and Distribution Learning
用于信号处理和分布学习的矩不变数据聚合
批准号:
2309570
负责人:
Anna Little
金额:
$36.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-06-01 至 2026-05-31

项目摘要

项目成果

Anna Little的其他基金

相似基金

相关文献

中文摘要
翻译
现代来自不同来源的数据爆炸式增长,极大地增加了对可靠数据集成工具的需求。在许多应用程序中,人们可以访问海量数据,但这些数据非常嘈杂,只有在聚合时才有意义。例如,在低温电子显微镜中,人们可以获得大量的分子图像。尽管如此,每一幅图像都有很高的噪声,并反映了随机的移动、方向和投影,这使得集成具有极大的挑战性。另一个例子是微阵列表达数据,通常是以小批量收集的,具有显著的批次效应。传统的统计方法是在整合前对数据进行标准化,即用每批数据的标准差减去平均值和标度,然后将标准化批次合并。然而,当每个批次/子总体的样本量很小时,这种方法将失败,因为批内均值和方差的估计可能不可靠。该项目将开发计算高效和可伸缩的方法,用于从多个来源聚合噪声数据。该方法将应用于刑事司法系统的数据,以深入了解种族和基于性别的差异。该项目将支持几个层次的导师以及研究生和本科生的研究,并将贡献公开可用的代码。此外,这项研究还将成为信号处理和统计领域之间的桥梁。该项目的目的是开发数学工具,用于在两个不同的背景下对数据的第一和第二时刻不变地进行数据聚合。第一个背景是多参考比对(MRA)的数据聚合,这是一个由低温电子显微镜等生物学应用推动的主题。在经典的磁共振成像中,人们试图从许多噪声观测中恢复隐藏的信号,其中每个噪声观测都被随机转换并被加性噪声破坏。研究人员将探索这一模型的一般化,其中每个噪声观测也被随机的尺度变化所破坏。这项研究将开发一种计算高效的全信号恢复方法,该方法利用傅立叶和基于小波的特征来消除随机尺度变化对噪声数据的偏差。第二个背景是用于分布式学习的数据聚合。研究人员考虑了这样一种情况,即从不同的子总体收集数据以产生数据批次,但批次的样本量很小。在一个简单但令人信服的模型下,局部化因素只影响子总体的一阶和二阶矩,因此每批都由来自一些平移和重新缩放的普遍分布函数的独立、同分布的观测值组成。目标是通过聚合稀疏数据来恢复底层分布。这是一个高度相关的问题,因为它允许在传统方法失败的情况下对密度进行可靠的非参数估计,更广泛地说,它导致在数据很少的情况下对子总体进行更精确的比较。通过将这个问题视为一个广义的MRA问题,其中随机采样的不确定性取代了高斯噪声,该项目在信号处理方法的启发下为其解决方案贡献了一套创新的工具。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The explosion of data in the modern era from diverse sources has greatly increased the need for reliable data integration tools. In many applications, one has access to vast amounts of data, but the data is highly noisy and only meaningful when aggregated. For example, in cryo-electron microscopy, one has access to a large volume of molecular images. Still, each image is highly noisy and reflects a random shift, orientation, and projection, making integration highly challenging. Another example is microarray expression data, which is generally collected in small batches with significant batch effects. The traditional statistical approach is to “standardize the data” before integration, i.e., one subtracts the mean and scales by the standard deviation in each batch and then combines the standardized batches. However, when the sample size of each batch/subpopulation is small, this approach will fail since within-batch estimates of mean and variance can be unreliable. This project will develop computationally efficient and scalable methods for aggregating noisy data from multiple sources. The methodology will be applied to data from the criminal justice system to gain insights into race and gender-based discrepancies. The project will support several tiers of mentorship as well as graduate and undergraduate research and will contribute publicly available code. In addition, the research will also serve as a bridge between the fields of signal processing and statistics. The project aims to develop mathematical tools for data aggregation invariant to the first and second moments of the data in two distinct contexts. The first context is data aggregation for multi-reference alignment (MRA), a topic motivated by biological applications such as cryo-electron microscopy. In classic MRA one attempts to recover a hidden signal from many noisy observations, where each noisy observation has been randomly translated and corrupted by additive noise. The investigators will explore a generalization of this model where each noisy observation is also corrupted by a random scale change. The research will develop a computationally efficient method for full signal recovery that utilizes Fourier and wavelet-based features to unbias the noisy data for the random scale change. The second context is data aggregation for distribution learning. The investigators consider the scenario where data is collected from various sub-populations to produce data batches, but the sample sizes of the batches are small. Under a simple but compelling model, localization factors affect only the first and second moments of the sub-populations, so that each batch consists of independent, identically distributed observations from some shifted and rescaled universal distribution function. The goal is to recover the underlying distribution by aggregating the sparse data. This is a highly relevant problem, as it allows for reliable nonparametric estimates of the density in settings where traditional approaches fail and, more broadly, leads to more precise comparisons across sub-populations in settings where little data is available. By viewing this problem as a generalized MRA problem where uncertainty due to random sampling replaces Gaussian noise, this project contributes an innovative set of tools for its solution inspired by methods in signal processing.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Data-driven Path Metrics for Machine Learning
  • 批准号:
    2131292
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2021
  • 负责人:
    Anna Little
  • 依托单位:
Collaborative Research: Data-driven Path Metrics for Machine Learning
  • 批准号:
    1912906
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2019
  • 负责人:
    Anna Little
  • 依托单位:
海外基金