Moment Invariant Data Aggregation for Signal Processing and Distribution Learning
Moment Invariant Data Aggregation for Signal Processing and Distribution Learning
批准号:
2309570
负责人:
Anna Little
金额:
$36.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-06-01 至 2026-05-31
中文摘要
在现代时代,来自不同来源的数据爆炸式增长,极大地增加了对可靠数据集成工具的需求。在许多应用程序中,人们可以访问大量数据,但这些数据非常嘈杂,只有在聚合时才有意义。例如,在低温电子显微镜中,人们可以获得大量的分子图像。尽管如此,每个图像都是高噪声的,并且反映了随机的移位、方向和投影,使得集成具有高度挑战性。另一个例子是微阵列表达数据,它通常是以小批量收集的,具有显著的批量效应。传统的统计方法是在整合之前“标准化数据”,即,在每一批中减去平均值和标度乘以标准偏差,然后组合标准化的批次。然而,当每个批次/亚群的样本量很小时,这种方法将失败,因为平均值和方差的批内估计值可能不可靠。该项目将开发计算效率高且可扩展的方法,用于聚合来自多个来源的噪声数据。该方法将应用于刑事司法系统的数据,以深入了解基于种族和性别的差异。该项目将支持几个层次的导师以及研究生和本科生的研究,并将贡献公开可用的代码。此外,该研究还将成为信号处理和统计学领域之间的桥梁。该项目旨在开发数学工具,用于在两种不同的情况下对数据的一阶和二阶矩保持不变的数据聚合。第一个背景是多参考比对(MRA)的数据聚合,这是一个由冷冻电子显微镜等生物应用激发的主题。在经典的MRA中,人们试图从许多噪声观测中恢复隐藏的信号,其中每个噪声观测都被随机平移并被加性噪声破坏。研究人员将探索这个模型的推广,其中每个噪声观测也被随机尺度变化破坏。这项研究将开发一种计算效率高的方法,利用傅立叶和小波为基础的功能,以消除随机尺度变化的噪声数据的完整信号恢复。第二个背景是分布式学习的数据聚合。研究者考虑从不同亚群收集数据以产生数据批次的情况,但批次的样本量很小。在一个简单但令人信服的模型下,局部化因子只影响子总体的一阶矩和二阶矩,因此每一批都由来自一些移位和重新标度的普适分布函数的独立、同分布的观测值组成。目标是通过聚集稀疏数据来恢复底层分布。这是一个高度相关的问题,因为它允许在传统方法失败的情况下对密度进行可靠的非参数估计,并且更广泛地说,在数据很少的情况下,可以对各亚群进行更精确的比较。通过将该问题视为一个广义MRA问题,其中随机采样的不确定性取代了高斯噪声,该项目为信号处理方法的启发提供了一套创新的解决方案工具。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The explosion of data in the modern era from diverse sources has greatly increased the need for reliable data integration tools. In many applications, one has access to vast amounts of data, but the data is highly noisy and only meaningful when aggregated. For example, in cryo-electron microscopy, one has access to a large volume of molecular images. Still, each image is highly noisy and reflects a random shift, orientation, and projection, making integration highly challenging. Another example is microarray expression data, which is generally collected in small batches with significant batch effects. The traditional statistical approach is to “standardize the data” before integration, i.e., one subtracts the mean and scales by the standard deviation in each batch and then combines the standardized batches. However, when the sample size of each batch/subpopulation is small, this approach will fail since within-batch estimates of mean and variance can be unreliable. This project will develop computationally efficient and scalable methods for aggregating noisy data from multiple sources. The methodology will be applied to data from the criminal justice system to gain insights into race and gender-based discrepancies. The project will support several tiers of mentorship as well as graduate and undergraduate research and will contribute publicly available code. In addition, the research will also serve as a bridge between the fields of signal processing and statistics. The project aims to develop mathematical tools for data aggregation invariant to the first and second moments of the data in two distinct contexts. The first context is data aggregation for multi-reference alignment (MRA), a topic motivated by biological applications such as cryo-electron microscopy. In classic MRA one attempts to recover a hidden signal from many noisy observations, where each noisy observation has been randomly translated and corrupted by additive noise. The investigators will explore a generalization of this model where each noisy observation is also corrupted by a random scale change. The research will develop a computationally efficient method for full signal recovery that utilizes Fourier and wavelet-based features to unbias the noisy data for the random scale change. The second context is data aggregation for distribution learning. The investigators consider the scenario where data is collected from various sub-populations to produce data batches, but the sample sizes of the batches are small. Under a simple but compelling model, localization factors affect only the first and second moments of the sub-populations, so that each batch consists of independent, identically distributed observations from some shifted and rescaled universal distribution function. The goal is to recover the underlying distribution by aggregating the sparse data. This is a highly relevant problem, as it allows for reliable nonparametric estimates of the density in settings where traditional approaches fail and, more broadly, leads to more precise comparisons across sub-populations in settings where little data is available. By viewing this problem as a generalized MRA problem where uncertainty due to random sampling replaces Gaussian noise, this project contributes an innovative set of tools for its solution inspired by methods in signal processing.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Data-driven Path Metrics for Machine Learning
-
批准号:2131292
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2021
-
负责人:Anna Little
-
依托单位:
Collaborative Research: Data-driven Path Metrics for Machine Learning
-
批准号:1912906
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2019
-
负责人:Anna Little
-
依托单位:
海外基金