Reproducibility Assessment for Multivariate Assays
Reproducibility Assessment for Multivariate Assays
批准号:
8647816
负责人:
Chris Fraley
金额:
$13.11万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-04-01 至 2015-09-30
关键词:
AddressAffectAlgorithmsAreaBioinformaticsBiological AssayBiological FactorsBiological MarkersBiologyChIP-seqClinicalCloud ComputingDataData AnalysesDecision TreesDevelopmentDiagnosticDimensionsEffectivenessEvaluationEvolutionGenomicsGoalsGuidelinesIn VitroInvestigationLassoLeadLiteratureMachine LearningMeasurementMeasuresMedical ResearchMethodologyMethodsModelingMonitorOutcomePerformancePhasePlayProtocols documentationPublic HealthPublishingROC CurveReproducibilityResearch Project GrantsSchemeServicesSignal TransductionSimulateSmall Business Innovation Research GrantSourceSpecific qualifier valueStagingStatistical MethodsStatistical ModelsSupport SystemTechniquesTechnologyTherapeuticTreesValidationanalytical toolbaseclinical practicecostdata miningdesigndisease diagnosisdrug developmentfollow-upforesthigh throughput technologyimprovedindexingnovel diagnosticspublic health relevanceresearch study
中文摘要
项目摘要。该小型企业创新研究项目解决了评估的问题
分析高通量数据时的重复性。在具有大量有限元分析的数据的特征选择中
然而,众所周知,某些特征似乎会偶然影响结果,随后
基于这些特征的预测可能并不像最初的结果所显示的那样成功。
同样,在多变量分析中往往涉及多个阶段和许多参数。
签约分析高吞吐量配置文件。例如,一个特定的组合取得了良好的效果-
交叉验证实例的设置不能推广到其他实例。的目标是
这项提议是将评估重复实验重复性的新统计方法扩展到
机器学习的背景,并在此应用程序中展示有效性。机器学习
调查方法将包括随机森林、监督主成分、套索惩罚-
归一化和支持向量机。我们将使用来自基因组应用的模拟和真实数据来
展示此方法的潜力,以提供不与以下各项混淆的重复性评估
用于确定生物相关阈值的预先指定的选择,用于提高信号的准确性
识别,并用于识别次优结果。
关联性。尽管今天的高通量技术提供了革命性的临床
实践中,可用于从如此海量的数据中提取信息的分析工具尚不存在
已经充分发展,可以有针对性地探索潜在的生物学。该项目直接解决了
需要使FDA所称的IVDMIA(体外诊断多变量指数分析)透明,
可解释、可重现,因此是改进分析产品和服务的机会
提供给识别、表征和验证临床诊断和生物标记物的公司
药物开发决策点。拟议项目的长期目标是为
生物标记物发现和综合基因组分析,并将重复性评估纳入
多变量分析。这将使评估和改进检测生物病毒的方法成为可能
影响特定结果并导致更有效和更有效的疾病治疗方法的因素
诊断、治疗监测和治疗药物开发。
英文摘要
Project Summary. This Small Business Innovation Research project addresses the problem of assessing
reproducibility in analyzing high-throughput data. In feature selection for data with large numbers of fea-
tures, it is well known that some features will appear to affect an outcome by chance, and that subsequent
predictions based on these features may not be as successful as initial results would seem to indicate.
Similarly, there are often multiple stages, and many parameters, involved in the multivariate assays de-
signed to analyze high-throughput profiles. For example, good results achieved with a particular combina-
tion of settings for an instance of cross-validation may not generalize to other instances. The objective of
this proposal is to extend new statistical methods for assessing reproducibility in replicate experiments to
the context of machine learning, and demonstrate effectiveness in this application. The machine-learning
methods to be investigated will include random forests, supervised principal components, lasso penal-
ization and support vector machines. We will use simulated and real data from genomic applications to
show the potential of this approach for providing reproducibility assessments that are not confounded with
prespecified choices, for determining biologically relevant thresholds, for improving the accuracy of signal
identification, and for identifying suboptimal results.
Relevance. Although today's high-throughput technologies offer the possibility of revolutionizing clinical
practice, the analytical tools available for extracting information from this amount of data are not yet
sufficiently developed for targeted exploration of the underlying biology. This project directly addresses the
need to make what the FDA terms IVDMIA (In-Vitro Diagnostic Multivariate Index Assays) transparent,
interpretable, and reproducible, and is thus an opportunity to improve analysis products and services
provided to companies that identify, characterize, and validate biomarkers for clinical diagnostics and
drug development decision points. The long-term goal of the proposed project is to develop a platform for
biomarker discovery and integrative genomic analysis, with reproducibility assessment incorporated into
multivariate assays. This will enable evaluation and improvement of approaches to detecting the biological
factors that affect a particular outcome, and lead to more efficient and more effective methods for disease
diagnosis, treatment monitoring, and therapeutic drug development.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Parsimonious Models for Survival Data
-
批准号:8394875
-
项目类别:
-
资助金额:$15.0万
-
财政年份:2012
-
负责人:Chris Fraley
-
依托单位:
Parsimonious Models for Survival Data
-
批准号:8545192
-
项目类别:
-
资助金额:$7.28万
-
财政年份:2012
-
负责人:Chris Fraley
-
依托单位:
Least Angle Regression
-
批准号:7748342
-
项目类别:
-
资助金额:$16.82万
-
财政年份:2005
-
负责人:Chris Fraley
-
依托单位:
Least Angle Regression
-
批准号:7293630
-
项目类别:
-
资助金额:$15.85万
-
财政年份:2005
-
负责人:Chris Fraley
-
依托单位:
Software for Fitting Non-Gaussian Random Effects Models
-
批准号:7003818
-
项目类别:
-
资助金额:$37.85万
-
财政年份:2004
-
负责人:Chris Fraley
-
依托单位:
海外基金