xARA: ARA through Explainable AI
xARA: ARA through Explainable AI
批准号:
10547257
负责人:
Joseph Gormley
金额:
$65.58万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-01-24 至 2022-11-30
关键词:
中文摘要
为了响应NIH FOA OTA-19009“生物医学翻译:发展”,我们建议
建立一个自治中继代理(ARA),可以表征和评价质量
从多个多尺度异构知识提供者(KPs)返回的信息。
生物医学研究人员通过以下方式与知识提供者(KP)建立信任关系:
频繁和持续使用。随着时间的推移,一种熟悉感的发展推动了他们的理解,
洞察1)如何构建和调用更有效的查询,2)结果的质量,
可以期望响应于不同的查询参数和特征值,以及3)如何评估
特定查询结果的相关性。
虽然这种信息检索范式已经适度地服务于研究界
在过去,它是不可扩展的,KPs的数量、范围和复杂性都在以一种
(截至2019年1月报告的1,613个分子生物学数据库)。在这个曾经
随着信息格局的不断变化,生物医学研究人员现在有两种选择--
继续使用他们已经学会信任的少数几个关键点,但在可操作性方面仍然有限
他们将收到的信息,或投入时间,并接受使用一系列新的
信息资源很少或根本不熟悉,因此有效性不确定。如果研究人员
现在将从NIH和行业赞助的大量信息资产中受益
现有的和不断扩大的新的信息检索和质量评估技术将是
必需的.
我们建议建立一个解释性的自治中继代理(xARA),可以表征
通过对从多尺度异构KP返回的信息的质量进行评级来查询结果。
xARA将利用多种信息检索和可解释的人工智能(xAI)
策略来执行跨多个异构KP的查询,并通过以下方式对结果进行排名:
质量和相关性,同时还确定和解释
数据库用于相同的查询响应。为了实现这一承诺,我们将利用基于案例的
用生物医学数据训练的推理和语言模型(即,BioBERT和定制
通过Reactome和UniProt的注释嵌入)允许新级别的查询分析
和评估。
我们的策略将允许1)信息缺口,以填补测试替代查询模式
产生不同的表面语法,但拥有语义相关和可操作的概念,
2)对于给定的查询特征值要识别的不一致,以及3)识别和
通过相似性度量消除或合并语义冗余的查询结果,
可解释AI(xAI)社区采用的基于案例的推理策略,
机器学习模型的行为和性能。
本文提出的xARA功能将基于Weber博士实验室开发的策略
对于信息检索,在推理时希望更大的透明度,
实验数据是我们的首要目标。我们的多机构团队由高级
研究人员和软件工程师在计算机和数据方面接受过正式培训,
科学、化学信息学、生物信息学、分子生物学和生物化学。
查询异构KP的固有风险包括存在不一致的
在独特的KP数据结构中使用相同的生物医学概念。手动工程可能是
克服这些障碍是必要的,但对最初的国家来说不会是一个重大挑战。
原型,因为只有两个记录良好的关键点正在评估。另一个值得注意的风险是
从UniProt和Reactome生成的词嵌入的质量可能不是
足够,需要进一步对生物医学文本(如PubMed)进行文本分析,这是可行的
在我们项目计划的时间范围内。
英文摘要
In response to the NIH FOA OTA-19009 “Biomedical Translator: Development” we propose to
build an Autonomous Relay Agent (ARA) that can characterize and rate the quality of
information returned from multiple multiscale heterogeneous knowledge providers (KPs).
Biomedical researchers develop a trust relationship with a knowledge provider (KP) through
frequent and continued use. Over time a familiarity develops that drives their understanding and
insight on 1) how to structure and invoke more effective queries, 2) the quality of the results they
may expect in response to different query parameters and feature values, and 3) how to assess
the relevancy of a specific query’s results.
Although this information retrieval paradigm has served the research community moderately
well in the past it is not scalable and the number, scope and complexity of KPs is increasing at a
dramatic pace (1,613 molecular biology databases reported as of Jan. 2019). Within this ever
changing information landscape, a biomedical researcher now has two choices -- either
continue using the few KPs they have learned to trust but remain limited in the actionable
information they will receive, or invest the time and accept the risk of using a range of new
information resources with little or no familiarity and thus uncertain effectiveness. If researchers
are to benefit from the vast array of NIH and industry sponsored information assets now
available and expanding new information retrieval and quality assessment technologies will be
required.
We propose to build an Explanatory Autonomous Relay Agent (xARA) that can characterize
query results by rating the quality of information returned from multi-scale heterogeneous KPs.
The xARA will utilize multiple information retrieval and explainable Artificial Intelligence (xAI)
strategies to perform queries across multiple heterogeneous KPs and rank their results by
quality and relevancy while also identifying and explaining any inconsistencies among
databases for the same query response. To deliver on this promise, we will utilize case-based
reasoning and language models trained with biomedical data (i.e., BioBERT and custom
annotation embeddings through Reactome and UniProt) permitting a new level of query profiling
and assessment.
Our strategies will permit 1) information gaps to be filled by testing alternative query patterns
that produce different surface syntax yet possess semantically related and actionable concepts,
2) inconsistencies to be identified for a given query feature value, and 3) the identification and
elimination or merging of semantically redundant query results via similarity metrics enriched by
case-based reasoning strategies employed in the explainable AI (xAI) community to identify
machine learning model behavior and performance.
The xARA capabilities proposed herein will be based on strategies developed in Dr. Weber’s lab
for information retrieval where the desire for greater transparency when reasoning over
experimental data is our primary aim. Our multi-institutional team is comprised of senior
researchers and software engineers formally trained and experienced in the computer and data
sciences, cheminformatics, bioinformatics, molecular biology, and biochemistry.
Inherent risks in querying heterogeneous KPs include the presence of inconsistent labeling of
the same biomedical concept within unique KP data structures. Manual engineering may be
necessary to overcome such hurdles, but will not be a significant challenge for the initial
prototype, since only two well documented KPs are being evaluated. Another noteworthy risk is
that the quality of word embeddings generated from UniProt and Reactome may not be
sufficient, requiring further textual analysis of biomedical text like PubMed, which is feasible
within the timeframe of our project plan.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
xARA: ARA through Explainable AI
-
批准号:10706760
-
项目类别:
-
资助金额:$16.39万
-
财政年份:2020
-
负责人:Joseph Gormley
-
依托单位:
xARA: ARA through Explainable AI
-
批准号:10057158
-
项目类别:
-
资助金额:$79.59万
-
财政年份:2020
-
负责人:Joseph Gormley
-
依托单位:
xARA: ARA through Explainable AI
-
批准号:10330631
-
项目类别:
-
资助金额:$73.65万
-
财政年份:2020
-
负责人:Joseph Gormley
-
依托单位:
海外基金