课题基金 / 基金详情

BIGDATA: F: Statistical Approaches to Big Data Analytics

BIGDATA: F: Statistical Approaches to Big Data Analytics
BIGDATA:F:大数据分析的统计方法
批准号:
1633074
负责人:
James Marron
金额:
$50.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2021-08-31

项目摘要

项目成果

James Marron的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The goals of this project include developing new Big Data analytical methods, providing an insightful understanding of their properties, and demonstrating major improvements over existing methods. While the driving application is cancer research, the lessons learned will be broadly applicable to a wide array of Big Data contexts. The major challenges addressed here include Data Integration, Data Heterogeneity and Parallelization. Data Integration is a recently understood need for combining widely differing types of measurements made on a common set of subjects. For example, in cancer research, common measurements in modern Big Data sets include gene expression, copy number, mutations, methylation and protein expression. The development of deep new statistical methods is proposed which focus on central scientific issues such as how the various measurements interact with each other, and simultaneously on which aspects operate in an independent manner. Data Heterogeneity addresses a different issue which is also critical in cancer research. In this case, current efforts to boost sample sizes (essential to deeper scientific insights) involve multiple laboratories combining their data. A whole new conceptual model for understanding the bias-oriented challenges presented by this scenario, plus the foundations for the development of new analytical methods that are robust against such effects, will be developed here. Parallelization is the computational concept of doing large scale numerical analysis through the simultaneous use of multiple computer processors. The proposed research will provide new foundational understanding of several important issues in this area.Data Integration will center on the Joint and Individual Variation Explained methodology. Early versions have already provided scientific insights not available from previous data analytic approaches. The basic idea will be first extended in the direction of more insightful groupings of data blocks, essential for understanding the full breadth of relationships between the available measurement types. The second extension will be in the direction of divergent groups of subjects, very important to the study of subtypes in cancer research and to the rest of precision medicine. In addition to new methodology, new methods of validation are proposed, and an asymptotic study of the properties will be conducted. The key new concept behind Data Heterogeneity is to replace the usual Gaussian conceptual model with a Gaussian mixture model, which makes intuitive sense but creates challenges, for example when using likelihood approaches as the mixture distributions are not an exponential family. An even bigger challenge is that mere scale issues usually entail that full estimation of the distributional parameters is completely intractable. Yet many standard statistical methods can be negatively impacted by such structure in data, so the invention of a new class of statistical methods that are robust against this effect, without requiring full parameter estimation, are proposed. Validation and development of mathematical statistical insights will again be an important part of the research. Parallelization is an essential component of all modern computing environments. The proposed research takes a Fiducial Inference viewpoint, which gives new insights into how the needed numerical calculations can be farmed out to a variety of processors, and then the results combined into a useful analysis for complicated statistical tasks including hypothesis testing and construction of confidence intervals.
期刊论文(23)
专著(0)
科研奖励(0)
会议论文
BFF: Bayesian, Fiducial, Frequentist Analysis of Age Effects in Daily Diary Data
BFF:每日日记数据中年龄影响的贝叶斯、基准、频率分析
DOI: 10.1093/geronb/gbz100
发表时间: 2019
期刊: The Journals of Gerontology: Series B
影响因子: --
作者: [Neupert, Shevaun D., Hannig, Jan, Ram, ed., Nilam]
通讯作者: Ram, ed., Nilam
Generalized fiducial inference for logistic graded response models
逻辑分级响应模型的广义基准推理
DOI: 10.1007/s11336
发表时间: 2017
期刊: Psychometrika
影响因子: 3
作者: [Liu, Y., Hannig, J.]
通讯作者: Hannig, J.
Comments on “A Gibbs Sampler for a Class of Random Convex Polytopes”
对“一类随机凸多面体的吉布斯采样器”的评论
DOI: 10.1080/01621459.2021.1950002
发表时间: 2021
期刊: Journal of the American Statistical Association
影响因子: 3.7
作者: [Hoffman, Kentaro, Hannig, Jan, Zhang, Kai]
通讯作者: Zhang, Kai
A note on optimal sampling strategy for structural variant detection using optical mapping
关于使用光学映射进行结构变异检测的最佳采样策略的说明
DOI: 10.1080/03610926.2020.1723638
发表时间: 2021
期刊: Communications in Statistics - Theory and Methods
影响因子: --
作者: [Li, Weiwei, Hannig, Jan, Jones, Corbin D.]
通讯作者: Jones, Corbin D.
19
    Data Integration Via Analysis of Subspaces (DIVAS)
    Collaborative Research: Tree Structured Object Oriented Data Analysis
    Collaborative Research: Statistical Learning and Object Oriented Data Analysis
    High Dimension - Low Sample Size Statistical Analysis
    海外基金