Consistent model selection in the p>>n setting
Consistent model selection in the p>>n setting
批准号:
8451848
负责人:
Valen EARL Johnson
金额:
$28.09万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-04-01 至 2015-03-31
关键词:
AddressAlgorithmsAreaBiochemical GeneticsBiologicalBiological MarkersClassificationClinicalClipColorectal CancerDataData SetDevelopmentDiseaseDisease AttributesDisease OutcomeEnsureGene ExpressionGenesGoalsLassoLinear ModelsLog-Linear ModelsMalignant NeoplasmsMeasurementMeasuresMedicalMedical ResearchMethodologyMethodsModelingMolecularOutcomePathway interactionsPatientsPlant RootsPlayProbabilityProceduresProcessPropertyResearch PersonnelRoleSample SizeSchemeSelection CriteriaSeriesSingle Nucleotide PolymorphismSpace ModelsSpeedStatistical ModelsTechniquesTechnologyTestingbasedensitydiscrete datagenome wide association studyhigh throughput screeninginnovationmalignant breast neoplasmpublic health relevanceresearch studyresponsesimulationstatisticssuccesstool
中文摘要
描述(由申请人提供):医学研究中最基本和最常见的统计问题之一是模型选择问题。模型选择是研究人员确定测量量之间的关系的过程;因此,它在基本上所有高通量筛选数据的分析中发挥着核心作用。模型选择程序是发现疾病与大量生化、遗传和药理学变量之间联系的主要分析机制。在这项应用中测试的基本假设是,一类新的模型选择程序可以用于有效地识别生物变量和疾病结果之间的关联,即使在潜在的生物相关因素比对每个变量的观察结果多得多的情况下也是如此。该项目的目标是开发这些变量选择程序,以便它们可以应用于高通量筛选数据,并将由此产生的方法应用于三个重要的应用领域。为实现这些目标,将实现以下具体目标。建议的模型选择过程的已知理论性质将扩展到可用生物测量比每个测量的观测值多得多的情况(即,pn设置)。将确定对结果变量的最终模型中可以包括的变量数量的限制,并将开发有效的数值算法,以便这些方法可以应用于实际的高通量筛选数据。新的模型选择程序将用于定义二进制分类算法,该算法可以从高维基因表达数据集预测临床结果。新的模型选择程序将用于在使用单核苷酸多态数据的全基因组关联研究中识别和分析与癌症和其他疾病相关的基因之间的相互作用。新的模型选择程序将用于分析高通量分子询问数据所提供的生物途径。在该项目期间开发的算法构成了模型选择领域的一项重大创新,并将为医学研究人员提供一套新的独特工具,用于从高通量筛查数据中有效地识别生物标记物、疾病属性和患者结果之间的生物关联。
英文摘要
DESCRIPTION (provided by applicant): Among the most fundamental and commonly encountered statistical problems in medical research is the problem of model selection. Model selection is the process by which researchers identify the relationships between measured quantities; thus it plays a central role in the analysis of essentially all high-throughput screening data. Model selection procedures represent the primary analytical mechanism through which the associations between diseases and large numbers of biochemical, genetic and pharmacological variables are discovered. The fundamental hypothesis tested in this application is that a new class of model selection procedures can be used to effectively identify associations between biological variables and disease outcomes, even in settings where there are many more potential biological correlates than there are observations on each variable. The goals of this project are to develop these variable selection procedures so that they can be applied to high-throughput screening data, and to apply the resulting methodology in three important application areas. To achieve these goals, the following specific aims will be addressed. Known theoretical properties of the proposed model selection procedures will be extended to cases in which there are many more biological measurements available than there are observations on each measurement (i.e., p n setting). Constraints on the number of variables that can be included in final models for outcome variables will be determined, and efficient numerical algorithms will be developed so that these methods can be applied to actual high-throughput screening data. The new model selection procedures will be used to define binary classification algorithms that can predict clinical outcomes from high-dimensional gene expression data sets. The new model selection procedures will be used to identify and analyze interactions between genes that are associated with cancer and other diseases in genome-wide association studies using single-nucleotide polymorphism data. The new model selection procedures will be used to analyze biological pathways as informed by high- throughput molecular interrogation data. The algorithms developed during this project constitute a major innovation in the field of model selection and will provide medical researchers with a new and unique set of tools for effectively identifying biological associations among biomarkers, disease attributes, and patient outcomes from high-throughput screening data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Consistent model selection in the p>>n setting
-
批准号:8084684
-
项目类别:
-
资助金额:$31.03万
-
财政年份:2011
-
负责人:Valen EARL Johnson
-
依托单位:
Consistent model selection in the p>>n setting
-
批准号:8235772
-
项目类别:
-
资助金额:$29.13万
-
财政年份:2011
-
负责人:Valen EARL Johnson
-
依托单位:
Consistent model selection in the p>>n setting
-
批准号:8646886
-
项目类别:
-
资助金额:$29.02万
-
财政年份:2011
-
负责人:Valen EARL Johnson
-
依托单位:
Consistent variable selection in p>>n settings
-
批准号:9106867
-
项目类别:
-
资助金额:$33.32万
-
财政年份:2011
-
负责人:Valen EARL Johnson
-
依托单位:
Consistent variable selection in p>>n settings
-
批准号:9340069
-
项目类别:
-
资助金额:$32.11万
-
财政年份:2011
-
负责人:Valen EARL Johnson
-
依托单位:
RECONSTRUCTION AND ANALYSIS OF EMISSION TOMOGRAPHY DATA
-
批准号:2097473
-
项目类别:
-
资助金额:$11.35万
-
财政年份:1992
-
负责人:Valen EARL Johnson
-
依托单位:
RECONSTRUCTION AND ANALYSIS OF EMISSION TOMOGRAPHY DATA
-
批准号:3460491
-
项目类别:
-
资助金额:$9.43万
-
财政年份:1992
-
负责人:Valen EARL Johnson
-
依托单位:
RECONSTRUCTION AND ANALYSIS OF EMISSION TOMOGRAPHY DATA
-
批准号:2097474
-
项目类别:
-
资助金额:$12.07万
-
财政年份:1992
-
负责人:Valen EARL Johnson
-
依托单位:
RECONSTRUCTION AND ANALYSIS OF EMISSION TOMOGRAPHY DATA
-
批准号:3460492
-
项目类别:
-
资助金额:$9.92万
-
财政年份:1992
-
负责人:Valen EARL Johnson
-
依托单位:
RECONSTRUCTION AND ANALYSIS OF EMISSION TOMOGRAPHY DATA
-
批准号:2097472
-
项目类别:
-
资助金额:$10.6万
-
财政年份:1992
-
负责人:Valen EARL Johnson
-
依托单位:
BIOSTATISTICAL AND DATABASE MANAGEMENT CORE
-
批准号:7928993
-
项目类别:
-
资助金额:$26.98万
-
财政年份:--
-
负责人:Valen EARL Johnson
-
依托单位:
BIOSTATISTICAL AND DATABASE MANAGEMENT CORE
-
批准号:8331189
-
项目类别:
-
资助金额:$24.91万
-
财政年份:--
-
负责人:Valen EARL Johnson
-
依托单位:
BIOSTATISTICAL AND DATABASE MANAGEMENT CORE
-
批准号:8133402
-
项目类别:
-
资助金额:$27.2万
-
财政年份:--
-
负责人:Valen EARL Johnson
-
依托单位:
BIOSTATISTICAL AND DATABASE MANAGEMENT CORE
-
批准号:8381068
-
项目类别:
-
资助金额:$26.38万
-
财政年份:--
-
负责人:Valen EARL Johnson
-
依托单位:
海外基金