Fully Nonparametric Models for Random Effects, Order Thresholding, Boostrap Testing, and Applications
Fully Nonparametric Models for Random Effects, Order Thresholding, Boostrap Testing, and Applications
批准号:
0805598
负责人:
Michael Akritas
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-07-15 至 2012-07-31
中文摘要
信息收集方面的技术进步,以及数学创新与生物、海洋/大气和社会心理科学的日益融合,在跨学科研究背景下创造了大量高度复杂和非常高维的数据集。这些数据的非标准特征包括非正态性、复杂的异质性和依赖结构、高维、小样本量和不平衡设计。研究者为这些数据集提出了更现实的统计模型,并开发了先进的统计分析方法。该项目有五个具体目标:通过提出交叉和嵌套双向随机和混合效应设计的完全非参数模型来增强建模替代方案,为每个模型下感兴趣的共同假设构建统计程序(包括稳健的基于秩的程序),提出顺序阈值(基于l统计的阈值)以降低替代假设的维数并识别信号位置。提出一种提高测试程序准确性的自举测试方法,并通过最近开发的基于测试的分类方法,探索上述测试程序在分类问题中的应用。所提出的模型和方法在方法上与标准似然方法、非参数和半参数模型以及贝叶斯技术有着根本的不同。这个项目的意义源于这样一个事实,即统计分析是许多昂贵的科学调查的最后阶段,而且往往是最重要的阶段。沿岸水域中某些污染物的浓度是否有下降或上升的趋势?污染物浓度的趋势是自然过程的结果还是由人类活动引起的?在不同的生物环境下,基因表达是否不同?哪些基因导致了这种差异?帮派抵抗教育和培训(G.R.E.A.T.)项目是否有效地减少了城市地区青少年的越轨行为/非法活动?在使用生物武器的早期检测中,是否有信号(某种高于背景率的症状),如果有,它位于何处?通常,为回答此类问题而收集的数据表现出高度非标准的特征。美国国家海洋和大气管理局监测海洋环境质量的国家现状和趋势项目的贻贝观察项目收集的数据是这类数据可以展示的复杂特征的一个很好的例子。由于其非标准特征,数据可能无法满足其他模型和方法所要求的规则性条件。在不满足的假设下分析数据可能会导致关于统计上重要的因素和趋势的错误结论,或者可能无法确定受疾病影响的现有信号或基因。本提案的目标是开发基于现实统计模型的先进数据分析方法,并开发实现这些方法的软件。
英文摘要
Technological advancements in information gathering, and the increased fusion of mathematical innovation with biological, oceanic/atmospheric, and psychosocial sciences, have created a plethora of highly complex and very high-dimensional data sets in interdisciplinary research contexts. The non-standard features of such data include non-normality, complex heterogeneity and dependence structures, high-dimensionality, low sample sizes and unbalanced designs. The investigator puts forth more realistic statistical models for such data sets and develops advanced statistical methods for their analysis. This project has five specific aims: to enhance the modeling alternatives by proposing fully nonparametric models for crossed and nested two-way random and mixed effects designs, to construct statistical procedures for the common hypotheses of interest under each of these models (including robust rank-based procedures), to propose order thresholding (thresholding based on L-statistics) for reducing the dimensionality of the alternative hypothesis and for identification of the signal location, to propose a bootstrap testing method for improved accuracy of the test procedures, and to explore applications of the aforementioned test procedures to classification problems, through the recently developed test-based classification method. The proposed models and methods are fundamentally different in approach from the standard likelihood methods, the non- and semi-parametric models, and the Bayesian techniques.The significance of this project stems from the fact that statistical analysis is the final, and often the most important, stage of many expensive scientific investigations. Does the concentration of certain contaminants in coastal waters have a decreasing or an increasing trend? Is a trend in the concentration of contaminants a result of natural processes or is it caused by human activity? Are gene expressions different under different biological environments and which genes are responsible for this difference? Has the Gang Resistance Education and Training(G.R.E.A.T.) program been effective in reducing adolescent deviant/illegal activities in urban areas? In early detection of the use of bioweapons, is there a signal (a certain symptom at rates higher than background) and if so where is it located? Typically, the data collected for answering such questions exhibit highly non-standard features. The data being collected by the Mussel Watch Project of NOAA's National Status and Trends program for monitoring marine environmental quality, is a good example of the type of complex features such data can exhibit. Due to their non-standard features, the data may fail to satisfy the regularity conditions that alternative models and methods require. Analyzing data under assumptions that are not satisfied may lead to incorrect conclusions regarding the statistically significant factors and trends, or may fail to identify an existing signal or the genes that are affected by a disease. The objective of this proposal is to develop advanced methods for data analysis based on realistic statistical models, and to develop software for their implementation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Variable Selection, Variable Screening and Dimension Reduction
-
批准号:1209059
-
项目类别:Continuing Grant
-
资助金额:$14.0万
-
财政年份:2012
-
负责人:Michael Akritas
-
依托单位:
Nonparametric Models and Methods for Social Sciences Data
-
批准号:0318200
-
项目类别:Standard Grant
-
资助金额:$23.42万
-
财政年份:2003
-
负责人:Michael Akritas
-
依托单位:
Collaborative Research: Nonparametric Models for Incomplete Clustered Data with Applications to the Social Sciences
-
批准号:9986592
-
项目类别:Continuing Grant
-
资助金额:$7.05万
-
财政年份:2000
-
负责人:Michael Akritas
-
依托单位:
Nonparametric Models and Methods for Analysis of Covariance in Social Sciences Research
-
批准号:9709891
-
项目类别:Continuing Grant
-
资助金额:$16.47万
-
财政年份:1997
-
负责人:Michael Akritas
-
依托单位:
Mathematical Sciences: Multivariate and Censored Data Analysis Methods for Astronomy
-
批准号:9208066
-
项目类别:Continuing Grant
-
资助金额:$16.0万
-
财政年份:1992
-
负责人:Michael Akritas
-
依托单位:
Mathematical Sciences: Advanced Statistical Methods for Analyzing Data from Astronomical Surveys
-
批准号:9007717
-
项目类别:Continuing Grant
-
资助金额:$6.85万
-
财政年份:1990
-
负责人:Michael Akritas
-
依托单位:
U.S.-Netherlands Cooperative Research: Statistical Methods for Analyzing Data Arising from Reliability Studies (Mathematical Sciences)
-
批准号:8700734
-
项目类别:Standard Grant
-
资助金额:$0.46万
-
财政年份:1987
-
负责人:Michael Akritas
-
依托单位:
海外基金