课题基金 / 基金详情

项目摘要

项目成果

LAWRENCE E HUNTER的其他基金

相似基金

相关文献

中文摘要
翻译
我们建议建立一个知识提供商,它将寻找、集成和提供人工智能就绪的、 通过生物医学文献的高性能文本挖掘的Biolink兼容模型。 翻译者目前挖掘我们想要的生物医学文献的问题 解决方案包括:(1)框架可扩展性和基准测试方面的缺陷 集成和验证新的文本挖掘方法很困难;(2)许可有问题 软件、术语和其他资源不能充分支持FIRE(和TLC) 最佳做法;(3)只处理PubMed的标题和摘要,不处理全文出版物;(4) 翻译人员使用较旧的NLP技术,性能相对较差;(5)缺乏 关于错误和其他问题的社区反馈机制;(6)缺乏连续性 更新以添加来自新出版物的知识;(7)输出知识表示,即 简单化和含糊,未能反映科学文献中所表达的内容的丰富性。 实施计划:我们的团队在NLP研究方面有很长的历史, 成功的开源软件项目、有效的基准和广泛的社区 订婚。我们将以NLM资助的信息提取工作的结果为基础,我们的 黄金标准的科罗拉多州注释丰富的全文(工艺)语料库,最近的BioNLP Open 我们组织的共享任务(BioNLP-OST)以及最先进的NLP的最新进展。 对于第1部分,我们将:(1)演示BioStack,一种可扩展的基于云的文本挖掘 生成基于开放生物医学本体的知识图的框架 (OBOS)。该BioStack演示将包括最先进的OBO概念识别器,可用于多个 本体,一种最先进的语义关系预测工具和一种最先进的 结构分析工具。所有生成的断言都将有起源元数据链接 对PMCID指定的文档中的特定文本范围的断言。(2)演示CRAFTST, 一种基于云的文本挖掘性能评价系统 系统与工艺黄金标准相抵触。(3)演示了一种自适应机器学习 说明如何高效地创建工具来提取Biolink关联类型的过程。 对于第二部分,我们建议扩展文本挖掘和评估框架,以使 通过Biolink和翻译社区,提高文本挖掘质量并扩展 挖掘的源文档的集合。具体地说,我们建议将10个长期目标 里程碑:(1)将飞船与Biolink对齐。(2)开发新的关联提取工具 文本。(3)开发和管理翻译者文本挖掘的社区参与过程。 (4)推广标杆管理。(5)提高召回率。(六)提高精准度。(七)提高计算量 效率。(8)扩展BioStack,将所有可用的全文生物医学期刊文章包括在内。(9) 扩大文件收集范围,包括专利和监管备案。(十)发展以科学家为本 改善对非开放出版商的文本挖掘的文档访问的运动。 生成的知识图可用于解决的问题类型包括 非常广泛,因为它是通过挖掘很大一部分生物医学文献产生的。 可以回答的问题包括关于特定断言的问题(例如,这种药物是不是 激动剂-激活剂?),一般关系(这两种蛋白质经常被提及吗 一起?)和文献(哪些出版物提到了这种基因、突变和药物?)。 整合:我们是开放科学界的长期贡献者,并拥有 与现有获奖者的长期合作;我们是NIH数据的参与者 下议院飞行员。我们建议通过以下方式使文本挖掘工具的输出与Biolink模型保持一致 OBO条款。我们建议在NIH云计算环境中实现我们的框架。 我们建议采用CD2H贡献者归因模型来前台社区 贡献。我们计划与NLM新生的基准活动和 SmartAPI努力构建转换器标准接口。 挑战和差距:高性能地从 文学仍然是一项艰巨的任务,不太可能在未来五年内得到解决。许多 由于许可证的限制,文本挖掘仍然无法访问重要的出版物。
英文摘要
We propose to build a knowledge provider that will seek out, integrate and provide AI-ready, BioLink-compatible models via high-performance text-mining of the biomedical literature. Problems with Translator’s current mining of the biomedical literature that we intend to solve include: (1) weaknesses in framework extensibility and benchmarking that make integrating and validating new text-mining approaches difficult; (2) problematic licensing of software, terminologies and other resources that do not adequately support FAIR (and TLC) best practices; (3) processing only PubMed titles and abstracts, not full text publications; (4) Translator’s use of older NLP technology with relatively poor performance; (5) lack of a mechanism for community feedback regarding errors and other problems; (6) lack of continuous updates to add knowledge from new publications; (7) output knowledge representation that is simplistic and vague, failing to reflect the richness of what is expressed in scientific documents. Plan for implementation: Our team has a long history of productive NLP research, successful open source software projects, effective benchmarking and broad community engagement. We will build on the results of NLM-funded work in information extraction, our gold-standard Colorado Richly Annotated Full Text (CRAFT) corpus, a recent BioNLP Open Shared Task (BioNLP-OST) that we organized, and recent advances in state-of-the-art NLP. For Segment 1, we will: (1) Demonstrate BioStacks, an extensible, cloud-based text-mining framework that produces knowledge graphs grounded in the Open Biomedical Ontologies (OBOs). This BioStacks demo will include a state-of-the-art OBO concept recognizer for multiple ontologies, a state-of-the-art semantic relationship prediction tool, and a state-of-the-art structural analysis tool. All generated assertions will have provenance metadata linking the assertion to a particular text span in a document specified by PMCID. (2) Demonstrate CRAFTST, a cloud-based text-mining evaluation system that evaluates the performance of text-mining systems against the CRAFT gold standard. (3) Demonstrate an adaptive machine learning process illustrating how to efficiently create tools to extract BioLink association types. For Segment 2, we propose to extend the text-mining and evaluation frameworks to align with BioLink and the Translator community, improve text-mining quality and expand the collection of source documents mined. Specifically, we propose to target 10 long term milestones: (1) Align CRAFT to BioLink. (2) Develop new tools for extracting associations from text. (3) Develop and manage a community engagement process on text-mining for Translator. (4) Extend benchmarking. (5) Improve recall. (6) Improve precision. (7) Improve computational efficiency. (8) Expand BioStacks to include all available full text biomedical journal articles. (9) Expand document collections to include Patents & Regulatory filings. (10) Develop a scientist-based movement to improve document access for text-mining from non-open publishers. The types of questions the resulting knowledge graph can be used to address are extremely broad, as it is generated by mining a large part of the biomedical literature. Questions that can be answered include those about specific assertions (e.g. is this drug an agonist-activator of this protein?), general relations (are these two proteins often mentioned together?), and documents (which publications mention this gene, mutation and drug?). Integration: We are long-time contributors to the open-science community and have longstanding collaborations with existing awardees; we were participants in the NIH Data Commons Pilot. We propose to align the output of text-mining tools to the BioLink model via OBO terms. We propose to implement our frameworks in NIH Cloud Computing environments. We propose to adopt the CD2H Contributor Attribution Model to foreground community contributions. We plan to coordinate with the NLM’s nascent benchmarking activities and the SmartAPI effort to build Translator standard interfaces. Challenges and gaps: High-performance mining of rich, contextualized knowledge from the literature remains a difficult task, and is unlikely to be solved in the next five years. Many important publications remain inaccessible to text-mining due to restrictive licensing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Scientific Questions: A New Target for Biomedical NLP
  • 批准号:
    10223438
  • 项目类别:
  • 资助金额:
    $45.31万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
Scientific Questions: A New Target for Biomedical NLP
  • 批准号:
    10454968
  • 项目类别:
  • 资助金额:
    $44.52万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
High Performance Text Mining for Translator
  • 批准号:
    10548337
  • 项目类别:
  • 资助金额:
    $46.61万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
Colorado Biomedical Informatics Training Program
  • 批准号:
    9526127
  • 项目类别:
  • 资助金额:
    $9.98万
  • 财政年份:
    2017
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
海外基金