课题基金 / 基金详情

Named Entity Recognition and Relationship Extraction in Biomedicine

Named Entity Recognition and Relationship Extraction in Biomedicine
生物医学中的命名实体识别和关系提取
批准号:
10261222
负责人:
Zhiyong Lu
金额:
$166.47万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:

项目摘要

项目成果

Zhiyong Lu的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Mining useful knowledge from the biomedical literature holds potentials for helping literature searching, automating biological data curation and many other scientific tasks. We have therefore focused on recognizing various types of biological entities in free text, such as gene/proteins, disease/conditions, and drug/chemicals, etc, and their relationships. Manually annotated data is key to developing text-mining and information-extraction algorithms. However, human annotation requires considerable time, effort and expertise. Given the rapid growth of biomedical literature, it is paramount to build tools that facilitate speed and maintain expert quality. While existing text annotation tools may provide user-friendly interfaces to domain experts, limited support is available for figure display, project management, and multi-user team annotation. In response, we developed TeamTat (https://www.teamtat.org), a web-based annotation tool (local setup available), equipped to manage team annotation projects engagingly and efficiently. TeamTat is a novel tool for managing multi-user, multi-label document annotation, reflecting the entire production life cycle. To facilitate research in the development of pre-training language representations in the biomedicine domain, we previously introduced the Biomedical Language Understanding Evaluation (BLUE) benchmark, upon which we recently studied a multi-task learning (MTL) model, as it has achieved remarkable success in natural language processing applications. Our empirical results demonstrate that the MTL fine-tuned models outperform state-of-the-art transformer models (eg, BERT and its variants) by 2.0% and 1.3% in biomedical and clinical domains, respectively. We make the datasets, pre-trained models, and codes publicly available. In 2019, we employed high-performance machine learning-based NER tools for concept recognition and trained our concept embeddings, BioConceptVec, via four different machine learning models on 30 million PubMed abstracts. BioConceptVec covers over 400,000 biomedical concepts mentioned in the literature and is of the largest among the publicly available biomedical concept embeddings to date. Our own research and work from others have shown that text-mining tools have rapidly matured: although not perfect, they now frequently provide outstanding results. In a recent community article, we describe 10 straightforward writing tipsand a web tool, PubReCheckguiding authors to help address the most common cases that remain difficult for text-mining tools. We anticipate these guides will help authors work be found more readily and used more widely, ultimately increasing the impact of their work and the overall benefit to both authors and readers. Separately, we have applied cutting-edge deep learning techniques to finding similar sentences in electric medical records, in a community task called Semantic Textual Similarity (BioCreative/OHNLP STS) challenge. The official results demonstrated our best submission was the ensemble of eight models. Our submissions achieved a Person correlation coefficient of 0.8328 the highest performance among all submission teams. As shown above, deep learning, a class of machine learning algorithms, has showed impressive results in several of our recent studies this year. In addition to applying it to natural language processing, we have also seen its success in our medical image analysis such as processing chest X-rays, CT images, and various kinds of retinal images for autonomous disease diagonosis and prognosis. As one of the most ubiquitous diagnostic imaging tests in medical practice, chest radiography requires timely reporting of potential findings and diagnosis of diseases in the images. Automated, fast, and reliable detection of diseases based on chest radiography is a critical step in radiology workflow. In this work, we developed and evaluated various deep convolutional neural networks (CNN) for differentiating between normal and abnormal frontal chest radiographs, in order to help alert radiologists and clinicians of potential abnormal findings as a means of work list triaging and reporting prioritization. A CNN-based model achieved an AUC of 0.98240.0043 (with an accuracy of 94.640.45%, a sensitivity of 96.500.36% and a specificity of 92.860.48%) for normal versus abnormal chest radiograph classification. The remarkable performance in diagnostic accuracy observed in this study shows that deep CNNs can accurately and effectively differentiate normal and abnormal chest radiographs, thereby providing potential benefits to radiology workflow and patient care. Another such project relates to Age-related macular degeneration (AMD), which is the leading cause of blindness in developed countries and, by 2040, will affect approximately 300 million people worldwide. Accurate AMD severity detection and progression prediction to sight-threatening late disease stage is thus of significant importance for personalizing monitoring and preventative interventions. As a joint effort between National Library of Medicine and National Eye Institute, we developed deep learning models for detecting reticular pseudodrusen (RPD) using fundus autofluorescence (FAF) images or, alternatively, color fundus photographs (CFP) in the context of age-related macular degeneration (AMD). We demonstrated that the fully automated deep learning model was superior to existing clinical standards and has the potential for improved patient care.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Named Entity Recognition and Relationship Extraction in Biomedicine
  • 批准号:
    9362446
  • 项目类别:
  • 资助金额:
    $140.39万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
  • 批准号:
    9564626
  • 项目类别:
  • 资助金额:
    $160.63万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Machine Learning and Natural Language Processing for Biomedical Applications
  • 批准号:
    10927050
  • 项目类别:
  • 资助金额:
    $387.34万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
  • 批准号:
    10007525
  • 项目类别:
  • 资助金额:
    $190.14万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
海外基金