课题基金 / 基金详情

PREMIERE: A PREdictive Model Index and Exchange REpository

PREMIERE: A PREdictive Model Index and Exchange REpository
PREMIERE:预测模型索引和交换存储库
批准号:
10228009
负责人:
ALEX BUI
金额:
$68.24万
依托单位国家:
美国
项目类别:
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-09-15 至 2024-05-31

项目摘要

项目成果

ALEX BUI的其他基金

相似基金

相关文献

中文摘要
翻译
融合新的机器学习(ML)数据驱动方法;提高计算能力;以及 访问丰富的电子健康记录(EHR)和其他紧急类型的数据(例如,组学、成像、 MHealth)正在加快生物医学预测模型的开发。这些模型的范围从传统的 统计方法(例如,回归)到更高级的深度学习技术(例如,卷积 神经网络),并跨越不同的任务(例如,生物标记物/途径发现、诊断、预后)。 两个问题已经变得明显:1)由于没有全面的标准来支持传播 考虑到解释和实施方面的挑战,这些模型的科学重现性是有问题的;以及 2)随着新模型的提出,评估业绩差异的方法以及对外部环境的洞察 有效性(即可运输性)是必要的。超越数据共享和建模“可执行文件”的工具 获取完全复制模型及其评估所需的(元)数据。 本R01的目标是开发一种信息学标准,以支持 统计和基于ML的生物医学预测模型的科学重复性;在此基础上,我们然后 开发新的计算方法来比较模型的性能。我们从扩展电流开始 预测模型标记语言(PMML)标准,用于全面表征生物医学数据集并协调 变量定义;阐明模型创建中涉及的算法(例如,数据预处理、参数 估计);并解释验证方法。重要的是,这种PMML格式的模型将成为 可查找、可访问、可互操作和可重复使用(即遵循公平原则)。然后我们提出了新的冰毒- 用于比较和对比预测模型,评估数据集之间的可传输性。当指标存在时 对于比较模型(例如,c-统计、校准),通常无法获得所需的病例级别信息 计算这些衡量标准。因此,我们介绍了一种基于模型报告的数据来模拟案例的方法。 Taset统计数据,支持这样的计算。然后将不同级别的可运输性分配给度量, 确定所选模型适用于给定人群/队列的程度(即,帮助回答 问题是,我可以将这个发布的模型与我自己的数据一起使用吗?)我们将这些努力结合在一起,在我们的建议中 框架,预测模型索引和交换储存库(首演)。我们将开发一个在线门户网站 和模型共享的资源库,我们的努力将包括培养一个用户社区 通过工作坊、模特儿和其他活动指导其发展。为了展示这些努力, 我们将使用来自目标领域的预测模型引导首映(基于成像的风险评估 肺癌筛查)。我们评估这些发展的努力将使一系列利益相关者参与(模型 开发人员、用户)告知我们标准的完整性;以及生物统计学家和临床专家指导 模型可运输性评估。
英文摘要
The confluence of new machine learning (ML) data-driven approaches; increased computational power; and access to the wealth of electronic health records (EHRs) and other emergent types of data (e.g., omics, imaging, mHealth) are accelerating the development of biomedical predictive models. Such models range from traditional statistical approaches (e.g., regression) through to more advanced deep learning techniques (e.g., convolutional neural networks, CNNs), and span different tasks (e.g., biomarker/pathway discovery, diagnostic, prognostic). Two issues have become evident: 1) as there are no comprehensive standards to support the dissemination of these models, scientific reproducibility is problematic, given challenges in interpretation and implementation; and 2) as new models are put forth, methods to assess differences in performance, as well as insights into external validity (i.e., transportability), are necessary. Tools moving beyond the sharing of data and model “executables” are needed, capturing the (meta)data necessary to fully reproduce a model and its evaluation. The objective of this R01 is the development of an informatics standard supporting the requisite information for scientific reproducibility for statistical and ML-based biomedical predictive models; from this foundation, we then develop new computational approaches to compare models' performance. We begin by extending the current Predictive Model Markup Language (PMML) standard to fully characterize biomedical datasets and harmonize variable definitions; to elucidate the algorithms involved in model creation (e.g., data preprocessing, parameter estimation); and to explain the validation methodology. Importantly, models in this PMML format will become findable, accessible, interoperable, and reusable (i.e., following FAIR principles). We then propose novel meth- ods to compare and contrast predictive models, assessing transportability across datasets. While metrics exist for comparing models (e.g., c-statistics, calibration), often the required case-level information is not available to calculate these measures. We thus introduce an approach to simulate cases based on a model's reported da- taset statistics, enabling such calculations. Different levels of transportability are then assigned to the metrics, determining the extent to which a selected model is applicable to a given population/cohort (i.e., helping answer the question, can I use this published model with my own data?). We tie these efforts together in our proposed framework, the PREdictive Model Index & Exchange REpository (PREMIERE). We will develop an online portal and repository for model sharing around PREMIERE, and our efforts will include fostering a community of users to guide its development through workshops, model-thons, and other activities. To demonstrate these efforts, we will bootstrap PREMIERE with predictive models from a targeted domain (risk assessment in imaging-based lung cancer screening). Our efforts to evaluate these developments will engage a range of stakeholders (model developers, users) to inform the completeness of our standard; and biostatisticians and clinical experts to guide assessment of model transportability.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Building BRIDGEs: Coordinating Standards, Diversity, and Ethics to Advance Biomedical AI
Building BRIDGEs: Coordinating Standards, Diversity, and Ethics to Advance Biomedical AI
Building BRIDGEs: Coordinating Standards, Diversity, and Ethics to Advance Biomedical AI
Predicting who will fracture: Exploration of machine learning in the observational Women's Health Initiative Study dataset.
海外基金