课题基金 / 基金详情

项目摘要

项目成果

LAWRENCE E HUNTER的其他基金

相似基金

相关文献

中文摘要
翻译
我们建议建立一个知识提供者,将寻找,整合和提供AI就绪, 通过生物医学文献的高性能文本挖掘的BioLink兼容模型。 译者目前对生物医学文献的挖掘存在的问题, 解决方案包括:(1)框架可扩展性和基准测试方面弱点, 整合和验证新的文本挖掘方法困难;(2)许可问题 不充分支持FAIR(和TLC)的软件、术语和其他资源 最佳实践;(3)仅处理PubMed标题和摘要,而不是全文出版物;(4) 翻译者使用较旧的NLP技术,性能相对较差;(5)缺乏 社区对错误和其他问题的反馈机制;(6)缺乏持续的 更新以添加来自新出版物的知识;(7)输出 简单和模糊,未能反映科学文献中表达的内容的丰富性。 实施计划:我们的团队在NLP研究方面有着悠久的历史, 成功的开源软件项目、有效的基准测试和广泛的社区 订婚我们将建立在NLM资助的信息提取工作的成果,我们的 黄金标准的科罗拉多丰富注释全文(CRAFT)语料库,最近的BioNLP开放 我们组织的共享任务(BioNLP-OST),以及最先进的NLP的最新进展。 对于第1部分,我们将:(1)演示BioStacks,一个可扩展的,基于云的文本挖掘 一个基于开放生物医学本体的知识图生成框架 (OBO)。这个BioStacks演示将包括一个最先进的OBO概念识别器, 本体,一个国家的最先进的语义关系预测工具,和一个国家的最先进的 结构分析工具。所有生成的断言都将具有出处元数据, 断言到由PMCID指定的文档中的特定文本范围。(2)展示手工艺, 基于云的文本挖掘评估系统,用于评估文本挖掘的性能 CRAFT黄金标准。(3)展示自适应机器学习 说明如何有效地创建提取BioLink关联类型的工具的过程。 对于Segment 2,我们建议扩展文本挖掘和评估框架, 与BioLink和Translator社区合作,提高文本挖掘质量, 收集挖掘的源文件。具体而言,我们建议以10个长期目标为目标, 里程碑:(1)将CRAFT与BioLink对齐。(2)开发新的工具,从 短信了(3)开发和管理Translator文本挖掘的社区参与流程。 (4)扩展基准测试。(5)提高记忆力。(6)提高精度。(7)提高计算 效率(8)扩展BioStacks以包括所有可用的全文生物医学期刊文章。(九) 扩大文件收集范围,包括专利和监管文件。(10)建立一个以科学家为基础的 这是一项旨在改善非开放出版商的文本挖掘文档访问的运动。 生成的知识图谱可用于解决的问题类型包括 它非常广泛,因为它是通过挖掘大部分生物医学文献而产生的。 可以回答的问题包括关于特定断言的问题(例如,这种药物是 这种蛋白质的激动剂-激活剂?),一般关系(这两种蛋白质经常被提及 在一起?),和文件(哪些出版物提到了该基因、突变和药物?)。 整合:我们是开放科学社区的长期贡献者, 与现有获奖者的长期合作;我们是NIH数据的参与者 Commons Pilot.我们建议通过以下方式将文本挖掘工具的输出与BioLink模型相匹配: OBO术语。我们建议在NIH云计算环境中实现我们的框架。 我们建议采用CD 2 H贡献者归因模型对前台社区进行分析 捐款.我们计划与NLM新生的基准测试活动和 SmartAPI致力于构建Translator标准接口。 挑战和差距:高效挖掘来自 文学仍然是一项艰巨的任务,在未来五年内不可能得到解决。许多 由于限制性许可,文本挖掘仍然无法访问重要的出版物。
英文摘要
We propose to build a knowledge provider that will seek out, integrate and provide AI-ready, BioLink-compatible models via high-performance text-mining of the biomedical literature. Problems with Translator’s current mining of the biomedical literature that we intend to solve include: (1) weaknesses in framework extensibility and benchmarking that make integrating and validating new text-mining approaches difficult; (2) problematic licensing of software, terminologies and other resources that do not adequately support FAIR (and TLC) best practices; (3) processing only PubMed titles and abstracts, not full text publications; (4) Translator’s use of older NLP technology with relatively poor performance; (5) lack of a mechanism for community feedback regarding errors and other problems; (6) lack of continuous updates to add knowledge from new publications; (7) output knowledge representation that is simplistic and vague, failing to reflect the richness of what is expressed in scientific documents. Plan for implementation: Our team has a long history of productive NLP research, successful open source software projects, effective benchmarking and broad community engagement. We will build on the results of NLM-funded work in information extraction, our gold-standard Colorado Richly Annotated Full Text (CRAFT) corpus, a recent BioNLP Open Shared Task (BioNLP-OST) that we organized, and recent advances in state-of-the-art NLP. For Segment 1, we will: (1) Demonstrate BioStacks, an extensible, cloud-based text-mining framework that produces knowledge graphs grounded in the Open Biomedical Ontologies (OBOs). This BioStacks demo will include a state-of-the-art OBO concept recognizer for multiple ontologies, a state-of-the-art semantic relationship prediction tool, and a state-of-the-art structural analysis tool. All generated assertions will have provenance metadata linking the assertion to a particular text span in a document specified by PMCID. (2) Demonstrate CRAFTST, a cloud-based text-mining evaluation system that evaluates the performance of text-mining systems against the CRAFT gold standard. (3) Demonstrate an adaptive machine learning process illustrating how to efficiently create tools to extract BioLink association types. For Segment 2, we propose to extend the text-mining and evaluation frameworks to align with BioLink and the Translator community, improve text-mining quality and expand the collection of source documents mined. Specifically, we propose to target 10 long term milestones: (1) Align CRAFT to BioLink. (2) Develop new tools for extracting associations from text. (3) Develop and manage a community engagement process on text-mining for Translator. (4) Extend benchmarking. (5) Improve recall. (6) Improve precision. (7) Improve computational efficiency. (8) Expand BioStacks to include all available full text biomedical journal articles. (9) Expand document collections to include Patents & Regulatory filings. (10) Develop a scientist-based movement to improve document access for text-mining from non-open publishers. The types of questions the resulting knowledge graph can be used to address are extremely broad, as it is generated by mining a large part of the biomedical literature. Questions that can be answered include those about specific assertions (e.g. is this drug an agonist-activator of this protein?), general relations (are these two proteins often mentioned together?), and documents (which publications mention this gene, mutation and drug?). Integration: We are long-time contributors to the open-science community and have longstanding collaborations with existing awardees; we were participants in the NIH Data Commons Pilot. We propose to align the output of text-mining tools to the BioLink model via OBO terms. We propose to implement our frameworks in NIH Cloud Computing environments. We propose to adopt the CD2H Contributor Attribution Model to foreground community contributions. We plan to coordinate with the NLM’s nascent benchmarking activities and the SmartAPI effort to build Translator standard interfaces. Challenges and gaps: High-performance mining of rich, contextualized knowledge from the literature remains a difficult task, and is unlikely to be solved in the next five years. Many important publications remain inaccessible to text-mining due to restrictive licensing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Scientific Questions: A New Target for Biomedical NLP
  • 批准号:
    10223438
  • 项目类别:
  • 资助金额:
    $45.31万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
Scientific Questions: A New Target for Biomedical NLP
  • 批准号:
    10454968
  • 项目类别:
  • 资助金额:
    $44.52万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
High Performance Text Mining for Translator
  • 批准号:
    10548337
  • 项目类别:
  • 资助金额:
    $46.61万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
Colorado Biomedical Informatics Training Program
  • 批准号:
    9526127
  • 项目类别:
  • 资助金额:
    $9.98万
  • 财政年份:
    2017
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
海外基金