课题基金 / 基金详情

项目摘要

项目成果

Armin Schwartzman的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):在使用高通量技术寻找疾病和健康风险标记物时,大规模多项检测已变得无处不在。虽然多个测试的统计方法通常假设测试之间是独立的,但许多实际情况显示出相关性和基本结构。空间结构的例子是蛋白质组数据的一维(1D);环境数据的2D;以及大脑成像数据的3D。在分析中忽略相关性可能导致发现的特征的不同集合和排序,从而导致增加的错误率和重要特征的潜在遗漏。有必要描述多重测试中相关性的影响,并将其纳入分析。这项建议的目标是开发多种测试方法,将相关性纳入数据,以增加统计能力,控制错误率,并获得适当的可解释结果。这是通过两种不同的方式完成的。(1)在目标1和目标2中,我们假设空间结构和平稳的遍历相关,其中感兴趣的信号由相对较少的单峰组成。我们使用随机场理论来计算p值,以检验平滑后观测数据的局部极大值的高度。我们将这些方法的复杂性从一维发展到三维域,从等宽的峰发展到不等宽的峰。然后,我们将这些方法应用于从高通量技术获得的各种类型的数据,特别是:用于识别癌症蛋白质生物标志物的质谱学数据;用于识别由于气候变化而面临热应激风险的地理区域的气候模型输出数据;以及用于识别涉及认知发育异常的解剖区域的大脑成像数据。(2)在目标3中,我们假设了一种一般的相关性结构,不一定是平稳的或遍历的,并提出了一种条件边缘分析,其中通过对观察到的可能零案例的边缘分布的条件来纳入相关性。虽然不是排他性的,但始终强调错误发现率推断。这一建议为随机领域的信号检测提供了一个统一的观点,该观点广泛适用于从蛋白质组学到医学成像到环境监测的一大类问题。从统计学的角度来看,它为随机场中的FDR控制问题提供了一个新的答案。通过利用相关性结构,本方案中开发的方法在寻找标记时提供了更高的统计能力,因此在后续研究中将测试较少数量的错误标记。
英文摘要
DESCRIPTION (provided by applicant): Large-scale multiple testing has become ubiquitous in the search for disease and health risk markers using high-throughput technologies. While statistical methods for multiple testing often assume independence between the tests, many real situations exhibit dependence and an underlying structure. Examples of spatial structure are one-dimensional (1D) in the case of proteomic data; 2D in the case of environmental data; and 3D in the case of brain imaging data. Ignoring correlation in the analysis may lead to a different set and ordering of discovered features, resulting in increased error rates and potential missing of important features. There is a need to characterize the effect of correlation in multiple testing and incorporate it into the analysis. The goal of this proposal is to develop multiple testing methods that incorporate the correlation in the data in order to increase statistical power, control error rates and obtain appropriately interpretable results. This is done in two different ways. (1) In Aims 1 and 2, we assume a spatial structure and stationary ergodic correlation, where the signal of interest consists of a relatively small number of unimodal peaks. We use random field theory to compute p-values for testing the heights of local maxima of the observed data after smoothing. We develop these methods in complexity from 1D to 3D domains, and from peaks of equal width to peaks of unequal width. We then adapt and apply these methods to various types of data obtained from high-throughput technologies, specifically: mass- spectrometry data for identifying protein biomarkers of cancer; climate model output data for identification of geographical regions at risk for heat stress as a result of climate change; and brain imaging data for identification of anatomical regions involved in abnormal cognitive development. (2) In Aim 3, we assume a general correlation structure, not necessarily stationary or ergodic, and propose a conditional marginal analysis, where correlation is incorporated through conditioning on the observed marginal distribution of likely null cases. Although not exclusively, emphasis throughout is placed on false discovery rate inference. This proposal provides a unified view of signal detection for random fields that applies broadly to a large class of problems ranging from proteomics to medical imaging to environmental monitoring. From a statistical point of view, it provides a new answer to the problem of controlling FDR in random fields. By taking advantage of the dependence structure, the methods developed in this proposal offer higher statistical power in the search for markers, so that a smaller number of false markers will be tested in follow-up studies.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Estimating The Fraction of Variance Explained by Genetics and Neuroanatomy in Neuropsychiatric Conditions
Estimating The Fraction of Variance Explained by Genetics and Neuroanatomy in Neuropsychiatric Conditions
Spatial inference methods for image analysis
Spatial inference methods for image analysis
海外基金