课题基金 / 基金详情

项目摘要

项目成果

Zhiyong Lu的其他基金

相似基金

相关文献

中文摘要
翻译
从生物医学文献中挖掘有用的知识有助于文献检索、自动化生物数据管理和许多其他科学任务。因此,能够在自由文本中识别各种类型的生物实体非常重要,例如基因/蛋白质、疾病/状况和药物/化学品等。事实上,我们之前的PubMed日志分析显示,人们搜索某些生物医学概念的频率高于其他概念,并且不同概念之间存在强烈的关联。例如,疾病名称通常与基因/蛋白质和药物名称同时出现。我们过去的研究主要集中在PubMed引文中的基因和物种识别。2011-2012年,我们在继续致力于提高基因名称识别的同时,也将注意力转向了疾病名称检测。像基因一样,疾病名称也是不规则和模糊的,很难通过简单的字典查找方法来识别,这对文本挖掘社区来说是一项有趣的任务。然而,由于缺乏足够的训练数据,没有太多的工作集中在疾病名称识别。为此,我们创建了一个大型疾病语料库,包含793篇PubMed摘要中的6900个疾病名称。我们的数据语料库由12名注释者(每个注释2人)组成的团队开发,包含PubMed摘要中每种疾病发生的丰富注释。此外,疾病名称分为四个不同的组:特定疾病、疾病类别、综合提及和疾病修饰。当用作训练最先进的机器学习算法的黄金标准数据时,我们的数据比现有的具有有限注释的数据具有更高的性能。这些特点使我们的疾病名称语料库成为从生物医学文本中挖掘疾病相关信息的宝贵资源。
英文摘要
Mining useful knowledge from the biomedical literature holds potentials for helping literature searching, automating biological data curation and many other scientific tasks. Hence, it is important to be able to recognize various types of biological entities in free text, such as gene/proteins, disease/conditions, and drug/chemicals, etc. Indeed, our previous PubMed log analysis revealed that people search certain biomedical concepts more often than others and that there exist strong associations between different concepts. For example, a disease name often co-occurs with gene/proteins and drug names. Our own research in the past has mostly focused on identifying genes and species in PubMed citations. In 2011-2012, while continuing our efforts in improving gene name recognition, we also turned our attention to disease name detection. Like genes, disease names are also irregular and ambiguous, making them difficult to be identified through simple dictionary look-up methods and an interesting task for the text-mining community. However, due to the lack of adequate training data, there has not been much work focused on disease name identification. To this end, we created a large-scale disease corpus consisting of 6,900 disease names in 793 PubMed abstracts. Developed by a team of 12 annotators (two people per annotation), our data corpus contains rich annotations for every disease occurrence in PubMed abstracts. Furthermore, disease names are categorized into four distinct groups: Specific Disease, Disease Class, Composite Mention and Disease Modifier. When used as the gold standard data for training state-of-the-art machine-learning algorithms, significantly higher performance was found on our data than an existing one with limited annotations. Such characteristics make our disease name corpus a valuable resource for mining disease-related information from biomedical text. Following named entity recognition, we also continued our research from previous years for automatically identifying relationships between various biological entities as an effort to build an end-to-end system that includes both entity recognition and relationship extraction. This year, our research emphasized on extracting pharmacogenomics (PGx) information from free text. Specifically, we developed a systematic approach to automatically identify PGx relationships between genes, drugs and diseases from trial records in ClinicalTrials.gov. In our evaluation, we found that our extracted relationships overlap significantly with the curated factual knowledge through the literature in a PGx database and that most relationships appear on average 5 years earlier in clinical trials than in their corresponding publications, suggesting that clinical trials may be valuable for both validating known and capturing new PGx related information in a more timely manner. Furthermore, two human reviewers judged a portion of computer-generated relationships and found an overall accuracy of 74% for our text-mining approach. This work has practical implications in enriching our existing knowledge on PGx gene-drug-disease relationships as well as suggesting crosslinks between ClinicalTrials.gov and other PGx knowledge bases. As mentioned earlier, one promising application area for text mining research is to assist manual literature curation, a highly time-consuming and labor-intensive process. In this regard, we conducted two separate investigations, one aiming to understand the needs of the curation community and the other directly improve links between literature and biological data. Together with colleagues outside of the NIH, we organized the BioCreative 2012 workshop on Interactive Text Mining in the Biocuration Workflow, an international event for bringing together the biocuration and text mining communities towards the development and evaluation of interactive text mining tools and systems to improve utility and usability in the biocuration workflow. Specifically, we chaired the Workshop Track II entitled Biocuration Workflows and Text Mining where we invited submissions of written descriptions of curation workflows from expert curated databases. We received seven qualified contributions, primarily from model organism databases such as FlyBase. Based on these descriptions, we identified commonalities and differences across the workflows, the common ontologies and controlled vocabularies used and the current and desired uses of text mining for biocuration. Compared to a similar study in 2009, our 2012 results show that many more databases are now using text mining in parts of their curation workflows. In addition, the Track II participants identified text-mining aids for finding gene names and symbols (gene indexing), prioritization of documents for curation (document triage), and ontology concept assignment as those most desired by the biocurators. Our second curation-oriented text mining research focused on directly improving links between literature and biological data. As we all know that in todays biomedical search, high-throughput experiments and bioinformatics techniques are creating an exploding volume of data that are becoming overwhelming to keep track of for biologists and researchers who need to access, analyze and process existing data. Much of the available data are being deposited in specialized databases, such as the Gene Expression Omnibus (GEO) for microarrays or the Protein Data Bank (PDB) for protein structures and coordinates. Data sets are also being described by their authors in publications archived in literature databases such as MEDLINE and PubMed Central. Currently, the curation of links between biological databases and the literature mainly relies on manual labor, which makes it a time-consuming and daunting task. Herein, we analyzed the current state of link curation between GEO, PDB and MEDLINE. We found that the link curation is heterogeneous depending on the sources and databases involved, and that overlap between sources is low, less than 50% for PDB and GEO. Furthermore, we showed that text-mining tools can automatically provide valuable evidence to help curators broaden the scope of articles and database entries that they review. As a result, we made recommendations to improve the coverage of curated links, as well as the consistency of information available from different databases while maintaining high-quality curation.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1016/j.jbi.2012.04.005
发表时间: 2012-10
期刊: JOURNAL OF BIOMEDICAL INFORMATICS
影响因子: 4.5
作者: [Li, Jiao, Lu, Zhiyong]
通讯作者: Lu, Zhiyong
A textual representation scheme for identifying clinical relationships in patient records.
用于识别患者记录中临床关系的文本表示方案。
DOI: 10.1109/icmla.2010.164
发表时间: 2011
期刊: Proceedings of the ... International Conference on Machine Learning and Applications. International Conference on Machine Learning and Applications
影响因子: --
作者: [Doğan,RezartaIslamaj, Névéol,Aurélie, Lu,Zhiyong]
通讯作者: Lu,Zhiyong
Author keywords in biomedical journal articles.
生物医学期刊文章中的作者关键词。
DOI: --
发表时间: 2010
期刊: AMIA ... Annual Symposium proceedings / AMIA Symposium. AMIA Symposium
影响因子: --
作者: [Neveol,Aurelie, Dogan,RezartaIslamaj, Lu,Zhiyong]
通讯作者: Lu,Zhiyong
DOI: 10.1186/1471-2105-12-s3-s3
发表时间: 2011-06-09
期刊: BMC bioinformatics
影响因子: 3
作者: [Islamaj Doğan R, Névéol A, Lu Z]
通讯作者: Lu Z
共 6 条
    Named Entity Recognition and Relationship Extraction in Biomedicine
    • 批准号:
      9362446
    • 项目类别:
    • 资助金额:
      $140.39万
    • 财政年份:
      --
    • 负责人:
      Zhiyong Lu
    • 依托单位:
    Query Log Analysis for Improving User Access to NCBI Web Services
    • 批准号:
      9564626
    • 项目类别:
    • 资助金额:
      $160.63万
    • 财政年份:
      --
    • 负责人:
      Zhiyong Lu
    • 依托单位:
    Machine Learning and Natural Language Processing for Biomedical Applications
    • 批准号:
      10927050
    • 项目类别:
    • 资助金额:
      $387.34万
    • 财政年份:
      --
    • 负责人:
      Zhiyong Lu
    • 依托单位:
    Named Entity Recognition and Relationship Extraction in Biomedicine
    • 批准号:
      10007525
    • 项目类别:
    • 资助金额:
      $190.14万
    • 财政年份:
      --
    • 负责人:
      Zhiyong Lu
    • 依托单位:
    海外基金