Named Entity Recognition and Relationship Extraction in Biomedicine
Named Entity Recognition and Relationship Extraction in Biomedicine
批准号:
8344935
负责人:
Zhiyong Lu
金额:
$49.97万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
AcetaminophenBenchmarkingBiologicalBiological databasesChemicalsClinicalCommunitiesDataData SetDatabasesDiseaseDoseEvaluationEventGene ProteinsGenesGoalsHealthInternationalKnowledgeLengthLiteratureMachine LearningMapsMeasuresMedicalMedical RecordsMethodsMiningModelingMonographNamesPatternPerformancePharmaceutical PreparationsPositioning AttributePubMedRelative (related person)ResearchSchemeSystemTechniquesTechnologyTextTylenolVocabularyWorkbaseheuristicsimprovedindexingopen sourcetext searchingtool
中文摘要
从生物医学文献中挖掘有用的知识对于促进文献检索、生物数据库管理和许多其他科学任务具有潜力。因此,重要的是能够识别自由文本中的各种类型的生物实体,例如基因/蛋白质,疾病/病症和药物/化学品等。事实上,我们以前的PubMed日志分析显示,人们搜索某些生物医学概念的频率高于其他概念,并且不同概念之间存在很强的关联。例如,在PubMed查询中,疾病名称通常与基因/蛋白质和药物名称共同出现。我们自己过去的研究集中在PubMed引文中识别基因和疾病。特别是,在2010年,我们共同组织了BioCreative III:一个国际挑战活动,让文本挖掘社区在PubMed Central的全长文章中寻找基因/蛋白质实体。
尽管有努力和进步,基因名称标准化(将基因名称映射到数据库标识符)仍然是一项具有挑战性的任务。部分原因是难以找到已识别的基因名称并将其与相应的物种相关联。这个问题的出现是因为物种信息往往没有明确说明旁边的基因/蛋白质提到或完全失踪的文章。因此,它需要自动的方法来推断这样的信息时,它是不容易获得的。为此,我们开发了一个名为SR 4GN的开源工具,用于在基因标准化的背景下进行物种识别和消歧。SR 4GN通过一组新的分类学显著扩展了我们以前的工作,用于识别文章中的焦点物种,并在无法找到此类信息时推断物种。根据我们对几个基准数据集的评估,SR 4GN达到了最先进的性能,并与其他类似系统相媲美。
今年关于实体识别的另一项研究是我们在PubMed Health药物专著中规范药物名称的工作。具体来说,我们开发了一个自动管道,用于根据自由文本中的成分和剂型(药物生产和分配的物理形式)在RxNorm(标准化药物词汇表)中识别药物概念。药物成分信息直接从专论标题中解析。对于剂型,开发了启发式规则和模式,以从全文各论正文中提取相关信息。与简单的查找方法相比,我们的方法在F-测度上显示出显著的改善。因此,这项研究被用来计算PubMed Health中每个药物专著的药物品牌名称列表。其结果已在PubMed Heath中部署和索引,以方便用户通过药物品牌访问相关药物页面(例如,搜索Tylenol以查看有关对乙酰氨基酚的信息)。
2011年,我们还探索了自动识别各种生物实体之间关系的方法,以努力构建一个包括实体识别和关系提取的端到端系统。在这项研究中,我们使用了来自第四次i2 b2挑战的数据,包括一个完全去识别的医疗记录语料库,其中包含临床概念(例如医疗问题)和关系(例如治疗改善医疗问题)的手动注释信息。机器学习是我们完成这项任务的主要方法。然而,与传统的词袋特征表示不同,我们表示了一种关系,该关系由文本中两个潜在相关概念的位置确定的五个不同的上下文块的方案:介绍性,第一概念,连接性,第二概念和结论性块。实验结果表明,当使用SVM时,这种新的上下文块表示优于传统的词袋模型。我们进一步的分析表明,这种表示的优点是它能够自动捕获概念之间的相对词阳性,这在其他研究中也很关键。
英文摘要
Mining useful knowledge from the biomedical literature holds potentials for facilitating literature search, biological database curation and many other scientific tasks. Hence, it is important to be able to recognize various types of biological entities in free text, such as gene/proteins, disease/conditions, and drug/chemicals, etc. Indeed, our previous PubMed log analysis revealed that people search certain biomedical concepts more often than others and that there exist strong associations between different concepts. For example, in PubMed queries a disease name often co-occurs with gene/proteins and drug names. Our own research in the past has focused on identifying genes and diseases in PubMed citations. In particular, in 2010 we co-organized BioCreative III: an international challenge event for engaging the text mining community on finding gene/protein entities in full-length articles from the PubMed Central.
Despite efforts and advances, gene name normalization (mapping a gene name to a database identifier) remains a challenging task. Partly, it is due to the difficulty in finding and associating the recognized gene name with its corresponding species. This problem arises because species information is often not explicitly stated next to the gene/protein mentions or completely missing in an article. Hence, it requires automatic methods to infer such information when it is not readily available. To this end, we have developed an open source tool called SR4GN for species recognition and disambiguation in the context of gene normalization. SR4GN significantly extends our previous work via a set of new heuristics for identifying focus species in an article and inferring species when such information cannot be found. According to our evaluation on several benchmark datasets, SR4GN achieves state-of-the-art performance and compares favorably to other similar systems.
Another research on entity recognition this year lies in our work on normalizing drug names in PubMed Health drug monographs. Specifically, we developed an automatic pipeline for identifying a drug concept in RxNorm (a standardized drug vocabulary) based on its ingredient and dose form (the physical form a drug is produced and dispensed) in free text. Drug ingredient information was directly parsed from the monograph title. As for the dose form, heuristic rules and patterns were developed to extract relevant information from the body of the full-text monographs. Compared with a simple lookup method, our method shows significant improvement in F-measure. As a result, this research is employed to compute a list of drug brand names for each drug monograph in PubMed Health. Its results have been deployed and indexed in PubMed Heath to facilitate user access to relevant drug pages through drug brands (e.g. searching Tylenol to see the information on Acetaminophen).
In 2011, we also explored means for automatically identifying relationships between various biological entities as an effort to build an end-to-end system that includes both entity recognition and relationship extraction. In this research, we used the data from the 4th i2b2 challenge comprising a corpus of fully de-identified medical records with manually annotated information for clinical concepts (e.g. medical problems) and relationships (e.g. treatments improve medical problems). Machine learning was our main approach for this task. However unlike the traditional bag-of-words feature representation, we represented a relationship with a scheme of five distinct context-blocks determined by the position of two potentially related concepts in the text: the introductory, first concept, connective, second concept, and conclusive block. Experimental results showed that when used with SVM, this new context-block representation outperformed the traditional bag-of-words model. Our further analysis suggested that the advantage of such a representation is its capability in automatically capturing the relative word positives between concepts, which has been found critical in other studies as well.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9362446
-
项目类别:
-
资助金额:$140.39万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:9564626
-
项目类别:
-
资助金额:$160.63万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Machine Learning and Natural Language Processing for Biomedical Applications
-
批准号:10927050
-
项目类别:
-
资助金额:$387.34万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:10007525
-
项目类别:
-
资助金额:$190.14万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
-
批准号:8149607
-
项目类别:
-
资助金额:$39.17万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9796762
-
项目类别:
-
资助金额:$225.49万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8558092
-
项目类别:
-
资助金额:$97.61万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8344934
-
项目类别:
-
资助金额:$49.97万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8943212
-
项目类别:
-
资助金额:$20.8万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8943240
-
项目类别:
-
资助金额:$83.19万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8558091
-
项目类别:
-
资助金额:$26.03万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:10261222
-
项目类别:
-
资助金额:$166.47万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9160930
-
项目类别:
-
资助金额:$40.9万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:10007518
-
项目类别:
-
资助金额:$213.91万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Machine learning for medical imaging: automated disease diagnosis and prognosis
-
批准号:10927041
-
项目类别:
-
资助金额:$138.33万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:10261212
-
项目类别:
-
资助金额:$170.11万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
国内基金
海外基金
企业绩效评价的DEA-Benchmarking方法及动态博弈研究
-
批准号:70571028
-
项目类别:面上项目
-
资助金额:16.5万元
-
批准年份:2005
-
负责人:杨印生
-
依托单位: