NSF Convergence Accelerator-Track D: Application of Sequential Inductive Transfer Learning for Experimental Metadata Normalization to Enable Rapid Integrative Analysis
NSF Convergence Accelerator-Track D: Application of Sequential Inductive Transfer Learning for Experimental Metadata Normalization to Enable Rapid Integrative Analysis
批准号:
2040521
负责人:
Grier Page
金额:
$100.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-09-15 至 2022-08-31
中文摘要
NSF融合加速器支持以使用为灵感,以团队为基础,多学科的努力,以应对国家重要性的挑战,并将在不久的将来为社会提供有价值的成果。这个项目,NSF融合加速器-轨道D:应用顺序归纳转移学习进行实验元数据规范化,以实现快速综合分析,开发工具来支持跨多个不同研究数据库的综合分析和Meta分析。随着数据驱动研究的爆炸式增长,研究人员甚至在一个单一的研究领域面临着许多不同的数据库,使用不同的术语和测量方案。因此,数据变得孤立-通过不同的过程收集,并由不同的元数据方案描述,没有数据库、元数据或变量的中央索引,使得研究人员难以识别用于综合分析或Meta分析的适当类型的数据。在这项工作的第一阶段,一个由统计学、流行病学、数据协调、机器学习、伦理学、数据库、成像和软件工程等领域的研究人员和专家组成的多学科团队将开发工具,将四个生物医学数据库的元数据链接起来,作为概念验证。链接的信息将通过该项目开发的MetaMatchMaker(3 M)门户提供。虽然传统的神经网络方法可用于链接实验元数据,但该方法可能很耗时,需要构建大型训练数据集。该项目采用了一种基于预训练学习模型(PLM)的替代方法,结合了自然语言处理(NLP)和迁移学习中使用的方法,允许将一个领域中构建的数据驱动模型应用于另一个领域,而无需开发大型训练数据集的时间和费用。在第一阶段,将从现有的PhenX-dbGAP元数据链接的大型手动训练数据集开发PLM,然后将用于链接来自四个不同生物医学数据库的元数据。第一阶段的结果将能够在更短的时间内快速和更广泛地识别实验数据,并减少用于数据标准化的资源;第二,PLM方法预计将通过消除或大大减少开发训练数据的需要,在跨数据库链接实验元数据方面节省大量资金。第二阶段的工作将扩大链接数据库的数量和种类,并使3 M公司符合为生物医学数据开发联合数据访问程序的要求,例如全球基因组学与健康联盟(GA 4GH)的身份验证和授权基础设施。这种方法成功的衡量标准包括提高速度和降低进行综合分析的成本;增加链接数据的重用。虽然第一阶段的概念验证是基于生物医学数据的链接,但如果成功,这种方法将适用于数据库frp,许多其他领域,包括国家安全,天气,环境研究,地球科学,天文学,法医分析,该奖项反映了NSF的法定使命,并已被认为是值得通过评估使用基金会的知识优点和更广泛的影响审查标准。
英文摘要
The NSF Convergence Accelerator supports use-inspired, team-based, multidisciplinary efforts that address challenges of national importance and will produce deliverables of value to society in the near future. This project, NSF Convergence Accelerator-Track D: Application of sequential inductive transfer learning for experimental metadata normalization to enable rapid integrative analysis, develops tools to support integrative analyses and meta analyses across multiple, distinct research databases. With the explosion in data-driven research, researchers even in a single domain of study are confronted with many different databases that use different terminologies and measurement schemes. Thus, data become siloed—collected via different processes and described by different metadata schemes with no central index of databases, metadata, or variables, making it difficult for a researcher to identify data of the appropriate type for use in integrative analyses or meta analyses. In Phase I of this effort, a multidisciplinary team of researchers and experts in statistics, epidemiology, data harmonization, machine leaning, ethics, databases, imaging, and software engineering will develop tools to link metadata across four biomedical database, as a proof of concept. The linked information will be available via the MetaMatchMaker (3M) portal to be developed by the project.While traditional neural network approaches could be used to link experimental metadata, that approach can be time consuming, requiring the construction of large training datasets. This project employs an alternative approach based on Pretrained Learning Models (PLMs), combining methods used in Natural Language Processing (NLP) and transfer learning, to allow for the application of data-driven models built in one domain to be applied to another, without the time and expense of developing large training datasets. In Phase I of the effort, a PLM will be developed from a large existing manually trained dataset of PhenX–dbGAP metadata linkage, which will then be used to link metadata from four diverse biomedical databases. The results from Phase I would enable rapid and broader identification of experimental data in less time and with fewer resources devoted to data normalization; second, the PLM approach is expected to provide significant savings in linking experimental metadata across databases by eliminating, or greatly reducing, the need for development of training data. Phase II of this effort will expand the number and variety of linked databases, and also make 3M compliant with developing federated data access procedures for biomedical data, such as Global Alliance for Genomics and Health (GA4GH)’s Authentication and Authorization Infrastructure. The metrics for success of this approach include increased speed and reduced cost of conducting integrative analyses; increased reuse of linked data. While, the proof of concept in Phase I is based on the linkage of biomedical data, if successful, this approach would be applicable to databases frp, many other domains including, for example, national security, weather, environmental research, geosciences, astronomy, forensic analysis, and law enforcement.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SGER: A Power and Sample Size Atlas for Microarray Research
-
批准号:0306596
-
项目类别:Standard Grant
-
资助金额:$9.78万
-
财政年份:2003
-
负责人:Grier Page
-
依托单位:
海外基金