Named Entity Recognition and Relationship Extraction in Biomedicine
Named Entity Recognition and Relationship Extraction in Biomedicine
批准号:
10261222
负责人:
Zhiyong Lu
金额:
$166.47万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
AddressAffectAge related macular degenerationAlgorithmsBenchmarkingBiologicalBlindnessChemicalsChestClassificationClinicalCodeColorCommunitiesDataData SetDetectionDeveloped CountriesDevelopmentDiagnostic ImagingDiseaseDrug InteractionsEvaluationFundusGene ProteinsGenesGoalsHumanImageImage AnalysisInformation RetrievalInternetJointsKnowledgeLabelLanguageLife Cycle StagesLiteratureMachine LearningManualsMedicalMedical ImagingMedical RecordsMethodsMiningModelingMonitorNamesNational Eye InstituteNatural Language ProcessingNetwork-basedOnline SystemsPatient CarePerformancePersonsPharmaceutical PreparationsPreventive InterventionProductionPubMedRadiology SpecialtyReaderReportingResearchSemanticsSeveritiesSpecificitySpeedTechniquesTechnologyTestingTextThoracic RadiographyTimeTrainingTriageUnited States National Library of MedicineVariantVisionWorkWritingX-Ray Computed Tomographyannotation systembaseconvolutional neural networkdata curationdeep learningdiagnostic accuracydisease diagnosisimprovedinterestlanguage trainingmachine learning algorithmmulti-task learningnoveloutcome forecastradiologistrapid growthresponseretinal imagingsuccesstext searchingtooluser-friendly
中文摘要
从生物医学文献中挖掘有用的知识有助于文献检索,自动化生物数据管理和许多其他科学任务。因此,我们专注于识别自由文本中的各种类型的生物实体,如基因/蛋白质,疾病/条件,药物/化学品等,以及它们之间的关系。
人工标注的数据是开发文本挖掘和信息提取算法的关键。然而,人工注释需要大量的时间、精力和专业知识。鉴于生物医学文献的快速增长,构建能够提高速度并保持专家质量的工具至关重要。虽然现有的文本注释工具可以为领域专家提供用户友好的界面,但对图形显示、项目管理和多用户团队注释的支持有限。作为回应,我们开发了TeamTat(https:www.teamtat.org),这是一个基于Web的注释工具(本地设置可用),可以有效地管理团队注释项目。TeamTat是一种新颖的工具,用于管理多用户、多标签文档注释,反映整个生产生命周期。
为了促进在生物医学领域开发预训练语言表示的研究,我们之前介绍了生物医学语言理解评估(BLUE)基准,我们最近研究了多任务学习(MTL)模型,因为它在自然语言处理应用中取得了显着的成功。我们的实证结果表明,MTL微调模型优于国家的最先进的Transformer模型(例如,BERT及其变体)的2.0%和1.3%,分别在生物医学和临床领域。我们将数据集、预训练模型和代码公开提供。
在2019年,我们采用了基于高性能机器学习的NER工具进行概念识别,并通过四种不同的机器学习模型在3000万份PubMed摘要上训练了我们的概念嵌入BioConceptVec。BioConceptVec涵盖了文献中提到的40多万个生物医学概念,是迄今为止公开的生物医学概念嵌入中最大的一个。
我们自己的研究和其他人的工作表明,文本挖掘工具已经迅速成熟:虽然不完美,但它们现在经常提供出色的结果。在最近的一篇社区文章中,我们描述了10个简单的写作技巧和一个网络工具,PubReCheck指导作者帮助解决文本挖掘工具仍然很难解决的最常见的情况。我们预计这些指南将帮助作者的作品更容易被发现,更广泛地使用,最终增加他们的工作的影响力和作者和读者的整体利益。
另外,我们还在一项名为语义文本相似性(BioCreative/OHNLP STS)挑战的社区任务中,应用了尖端的深度学习技术来寻找电子医疗记录中的相似句子。官方结果表明,我们最好的提交是八个模型的集合。我们提交的作品达到了0.8328的Person相关系数,是所有提交团队中表现最好的。
如上所示,深度学习是一类机器学习算法,在我们今年最近的几项研究中取得了令人印象深刻的结果。除了将其应用于自然语言处理之外,我们还在医学图像分析中看到了它的成功,例如处理胸部X射线,CT图像和各种视网膜图像,用于自主疾病诊断和预后。
作为医学实践中最普遍的诊断成像测试之一,胸部X射线摄影需要及时报告图像中的潜在发现和疾病诊断。基于胸部X线摄影的自动、快速和可靠的疾病检测是放射学工作流程中的关键步骤。在这项工作中,我们开发和评估了各种深度卷积神经网络(CNN),用于区分正常和异常的正面胸片,以帮助提醒放射科医生和临床医生潜在的异常发现,作为工作列表分类和报告优先级的一种手段。基于CNN的模型实现了正常与异常胸片分类的AUC为0.98240.0043(准确性为94.640.45%,灵敏度为96.500.36%,特异性为92.860.48%)。在这项研究中观察到的诊断准确性的显着表现表明,深CNN可以准确有效地区分正常和异常的胸片,从而为放射学工作流程和患者护理提供潜在的好处。
另一个这样的项目涉及视网膜相关性黄斑变性(AMD),这是发达国家失明的主要原因,到2040年,将影响全世界约3亿人。因此,准确的AMD严重程度检测和对威胁视力的晚期疾病阶段的进展预测对于个性化监测和预防性干预具有重要意义。作为国家医学图书馆和国家眼科研究所的共同努力,我们开发了深度学习模型,用于在年龄相关性黄斑变性(AMD)的背景下使用眼底自发荧光(FAF)图像或彩色眼底照片(CFP)检测网状假性玻璃疣(RPD)。我们证明了全自动深度学习模型上级现有的临床标准,并有可能改善患者护理。
英文摘要
Mining useful knowledge from the biomedical literature holds potentials for helping literature searching, automating biological data curation and many other scientific tasks. We have therefore focused on recognizing various types of biological entities in free text, such as gene/proteins, disease/conditions, and drug/chemicals, etc, and their relationships.
Manually annotated data is key to developing text-mining and information-extraction algorithms. However, human annotation requires considerable time, effort and expertise. Given the rapid growth of biomedical literature, it is paramount to build tools that facilitate speed and maintain expert quality. While existing text annotation tools may provide user-friendly interfaces to domain experts, limited support is available for figure display, project management, and multi-user team annotation. In response, we developed TeamTat (https://www.teamtat.org), a web-based annotation tool (local setup available), equipped to manage team annotation projects engagingly and efficiently. TeamTat is a novel tool for managing multi-user, multi-label document annotation, reflecting the entire production life cycle.
To facilitate research in the development of pre-training language representations in the biomedicine domain, we previously introduced the Biomedical Language Understanding Evaluation (BLUE) benchmark, upon which we recently studied a multi-task learning (MTL) model, as it has achieved remarkable success in natural language processing applications. Our empirical results demonstrate that the MTL fine-tuned models outperform state-of-the-art transformer models (eg, BERT and its variants) by 2.0% and 1.3% in biomedical and clinical domains, respectively. We make the datasets, pre-trained models, and codes publicly available.
In 2019, we employed high-performance machine learning-based NER tools for concept recognition and trained our concept embeddings, BioConceptVec, via four different machine learning models on 30 million PubMed abstracts. BioConceptVec covers over 400,000 biomedical concepts mentioned in the literature and is of the largest among the publicly available biomedical concept embeddings to date.
Our own research and work from others have shown that text-mining tools have rapidly matured: although not perfect, they now frequently provide outstanding results. In a recent community article, we describe 10 straightforward writing tipsand a web tool, PubReCheckguiding authors to help address the most common cases that remain difficult for text-mining tools. We anticipate these guides will help authors work be found more readily and used more widely, ultimately increasing the impact of their work and the overall benefit to both authors and readers.
Separately, we have applied cutting-edge deep learning techniques to finding similar sentences in electric medical records, in a community task called Semantic Textual Similarity (BioCreative/OHNLP STS) challenge. The official results demonstrated our best submission was the ensemble of eight models. Our submissions achieved a Person correlation coefficient of 0.8328 the highest performance among all submission teams.
As shown above, deep learning, a class of machine learning algorithms, has showed impressive results in several of our recent studies this year. In addition to applying it to natural language processing, we have also seen its success in our medical image analysis such as processing chest X-rays, CT images, and various kinds of retinal images for autonomous disease diagonosis and prognosis.
As one of the most ubiquitous diagnostic imaging tests in medical practice, chest radiography requires timely reporting of potential findings and diagnosis of diseases in the images. Automated, fast, and reliable detection of diseases based on chest radiography is a critical step in radiology workflow. In this work, we developed and evaluated various deep convolutional neural networks (CNN) for differentiating between normal and abnormal frontal chest radiographs, in order to help alert radiologists and clinicians of potential abnormal findings as a means of work list triaging and reporting prioritization. A CNN-based model achieved an AUC of 0.98240.0043 (with an accuracy of 94.640.45%, a sensitivity of 96.500.36% and a specificity of 92.860.48%) for normal versus abnormal chest radiograph classification. The remarkable performance in diagnostic accuracy observed in this study shows that deep CNNs can accurately and effectively differentiate normal and abnormal chest radiographs, thereby providing potential benefits to radiology workflow and patient care.
Another such project relates to Age-related macular degeneration (AMD), which is the leading cause of blindness in developed countries and, by 2040, will affect approximately 300 million people worldwide. Accurate AMD severity detection and progression prediction to sight-threatening late disease stage is thus of significant importance for personalizing monitoring and preventative interventions. As a joint effort between National Library of Medicine and National Eye Institute, we developed deep learning models for detecting reticular pseudodrusen (RPD) using fundus autofluorescence (FAF) images or, alternatively, color fundus photographs (CFP) in the context of age-related macular degeneration (AMD). We demonstrated that the fully automated deep learning model was superior to existing clinical standards and has the potential for improved patient care.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9362446
-
项目类别:
-
资助金额:$140.39万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:9564626
-
项目类别:
-
资助金额:$160.63万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Machine Learning and Natural Language Processing for Biomedical Applications
-
批准号:10927050
-
项目类别:
-
资助金额:$387.34万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:10007525
-
项目类别:
-
资助金额:$190.14万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
-
批准号:8149607
-
项目类别:
-
资助金额:$39.17万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8558092
-
项目类别:
-
资助金额:$97.61万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9796762
-
项目类别:
-
资助金额:$225.49万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8344934
-
项目类别:
-
资助金额:$49.97万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8943212
-
项目类别:
-
资助金额:$20.8万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8943240
-
项目类别:
-
资助金额:$83.19万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8558091
-
项目类别:
-
资助金额:$26.03万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9160930
-
项目类别:
-
资助金额:$40.9万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Machine learning for medical imaging: automated disease diagnosis and prognosis
-
批准号:10927041
-
项目类别:
-
资助金额:$138.33万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:10007518
-
项目类别:
-
资助金额:$213.91万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8344935
-
项目类别:
-
资助金额:$49.97万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:10261212
-
项目类别:
-
资助金额:$170.11万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
海外基金