课题基金 / 基金详情

BIGDATA: F: Statistical Approaches to Big Data Analytics

BIGDATA: F: Statistical Approaches to Big Data Analytics
BIGDATA:F:大数据分析的统计方法
批准号:
1633074
负责人:
James Marron
金额:
$50.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2021-08-31

项目摘要

项目成果

James Marron的其他基金

相似基金

相关文献

中文摘要
翻译
该项目的目标包括开发新的大数据分析方法,提供对其特性的深入了解,并展示对现有方法的重大改进。虽然推动应用的是癌症研究,但从中吸取的经验教训将广泛适用于广泛的大数据背景。这里涉及的主要挑战包括数据集成、数据异构性和并行化。数据集成是最近理解的一种需求,用于组合在一组公共对象上进行的广泛不同类型的测量。例如,在癌症研究中,现代大数据集中的常见衡量标准包括基因表达、拷贝数、突变、甲基化和蛋白质表达。提出了深层次的新统计方法的开发,这些方法侧重于核心的科学问题,如各种测量如何相互作用,以及同时关于哪些方面以独立的方式运作。数据异质性解决了一个不同的问题,这个问题在癌症研究中也是至关重要的。在这种情况下,目前扩大样本规模的努力(对更深入的科学洞察至关重要)涉及多个实验室合并他们的数据。这里将开发一个全新的概念模型,用于理解这一情景带来的偏向挑战,以及开发新的分析方法的基础,这些方法对这种影响具有强大的抵抗力。并行化是通过同时使用多个计算机处理器进行大规模数值分析的计算概念。拟议的研究将为这一领域的几个重要问题提供新的基础性理解。数据集成将以联合和个体差异解释方法为中心。早期版本已经提供了以前的数据分析方法所不具备的科学见解。基本思想将首先扩展到更有洞察力的数据块分组的方向,这对于理解可用测量类型之间的全面关系至关重要。第二个扩展将面向不同的受试组,这对癌症研究中的亚型研究和其他精确医学非常重要。除了新的方法外,还提出了新的验证方法,并将对这些性质进行渐近研究。数据异构性背后的关键新概念是用高斯混合模型取代通常的高斯概念模型,这在直觉上是有意义的,但也带来了挑战,例如在使用似然方法时,因为混合分布不是指数族。一个更大的挑战是,仅仅是规模问题通常会导致对分布参数的全面估计是完全难以处理的。然而,许多标准的统计方法可能会受到数据中这种结构的负面影响,因此提出了一种新的统计方法的发明,这些方法对这种影响具有鲁棒性,而不需要完整的参数估计。验证和发展数理统计洞察力将再次成为研究的重要部分。并行化是所有现代计算环境的重要组成部分。这项拟议的研究采用了Fiducial推理的观点,它为如何将所需的数值计算外包给各种处理器提供了新的见解,然后将结果合并为对复杂统计任务的有用分析,包括假设检验和构建可信区间。
英文摘要
The goals of this project include developing new Big Data analytical methods, providing an insightful understanding of their properties, and demonstrating major improvements over existing methods. While the driving application is cancer research, the lessons learned will be broadly applicable to a wide array of Big Data contexts. The major challenges addressed here include Data Integration, Data Heterogeneity and Parallelization. Data Integration is a recently understood need for combining widely differing types of measurements made on a common set of subjects. For example, in cancer research, common measurements in modern Big Data sets include gene expression, copy number, mutations, methylation and protein expression. The development of deep new statistical methods is proposed which focus on central scientific issues such as how the various measurements interact with each other, and simultaneously on which aspects operate in an independent manner. Data Heterogeneity addresses a different issue which is also critical in cancer research. In this case, current efforts to boost sample sizes (essential to deeper scientific insights) involve multiple laboratories combining their data. A whole new conceptual model for understanding the bias-oriented challenges presented by this scenario, plus the foundations for the development of new analytical methods that are robust against such effects, will be developed here. Parallelization is the computational concept of doing large scale numerical analysis through the simultaneous use of multiple computer processors. The proposed research will provide new foundational understanding of several important issues in this area.Data Integration will center on the Joint and Individual Variation Explained methodology. Early versions have already provided scientific insights not available from previous data analytic approaches. The basic idea will be first extended in the direction of more insightful groupings of data blocks, essential for understanding the full breadth of relationships between the available measurement types. The second extension will be in the direction of divergent groups of subjects, very important to the study of subtypes in cancer research and to the rest of precision medicine. In addition to new methodology, new methods of validation are proposed, and an asymptotic study of the properties will be conducted. The key new concept behind Data Heterogeneity is to replace the usual Gaussian conceptual model with a Gaussian mixture model, which makes intuitive sense but creates challenges, for example when using likelihood approaches as the mixture distributions are not an exponential family. An even bigger challenge is that mere scale issues usually entail that full estimation of the distributional parameters is completely intractable. Yet many standard statistical methods can be negatively impacted by such structure in data, so the invention of a new class of statistical methods that are robust against this effect, without requiring full parameter estimation, are proposed. Validation and development of mathematical statistical insights will again be an important part of the research. Parallelization is an essential component of all modern computing environments. The proposed research takes a Fiducial Inference viewpoint, which gives new insights into how the needed numerical calculations can be farmed out to a variety of processors, and then the results combined into a useful analysis for complicated statistical tasks including hypothesis testing and construction of confidence intervals.
期刊论文(23)
专著(0)
科研奖励(0)
会议论文
BFF: Bayesian, Fiducial, Frequentist Analysis of Age Effects in Daily Diary Data
BFF:每日日记数据中年龄影响的贝叶斯、基准、频率分析
DOI: 10.1093/geronb/gbz100
发表时间: 2019
期刊: The Journals of Gerontology: Series B
影响因子: --
作者: [Neupert, Shevaun D., Hannig, Jan, Ram, ed., Nilam]
通讯作者: Ram, ed., Nilam
Generalized fiducial inference for logistic graded response models
逻辑分级响应模型的广义基准推理
DOI: 10.1007/s11336
发表时间: 2017
期刊: Psychometrika
影响因子: 3
作者: [Liu, Y., Hannig, J.]
通讯作者: Hannig, J.
Comments on “A Gibbs Sampler for a Class of Random Convex Polytopes”
对“一类随机凸多面体的吉布斯采样器”的评论
DOI: 10.1080/01621459.2021.1950002
发表时间: 2021
期刊: Journal of the American Statistical Association
影响因子: 3.7
作者: [Hoffman, Kentaro, Hannig, Jan, Zhang, Kai]
通讯作者: Zhang, Kai
A note on optimal sampling strategy for structural variant detection using optical mapping
关于使用光学映射进行结构变异检测的最佳采样策略的说明
DOI: 10.1080/03610926.2020.1723638
发表时间: 2021
期刊: Communications in Statistics - Theory and Methods
影响因子: --
作者: [Li, Weiwei, Hannig, Jan, Jones, Corbin D.]
通讯作者: Jones, Corbin D.
共 19 条
    Data Integration Via Analysis of Subspaces (DIVAS)
    Collaborative Research: Tree Structured Object Oriented Data Analysis
    Collaborative Research: Statistical Learning and Object Oriented Data Analysis
    High Dimension - Low Sample Size Statistical Analysis
    海外基金