课题基金 / 基金详情

项目摘要

项目成果

Zhiyong Lu的其他基金

相似基金

相关文献

中文摘要
翻译
从生物医学文献中挖掘有用的知识有助于文献检索、自动化生物数据管理和许多其他科学任务。因此,能够在自由文本中识别各种类型的生物实体非常重要,例如基因/蛋白质、疾病/状况和药物/化学品等。事实上,我们之前的PubMed日志分析显示,人们搜索某些生物医学概念的频率高于其他概念,并且不同概念之间存在强烈的关联。例如,疾病名称通常与基因/蛋白质和药物名称同时出现。为了评估生物医学实体识别和关系提取的现状,我们在BioCreative V上组织了一场科学竞赛,这是一项评估生物学文本挖掘研究进展的国际挑战赛。具体来说,我们设计了两个挑战任务:疾病命名实体识别(DNER)和化学诱发疾病(CID)关系提取。为了帮助系统开发和评估,我们创建了一个大型的注释文本语料库,该语料库由1500篇PubMed文章中对化学物质、疾病及其相互作用的人类注释组成。全球34个团队参与了CDR任务:16个(DNER)和18个(CID)。最好的系统在DNER任务中获得了86.46%的f分,接近人类注释者间协议(0.8875),在CID任务中获得了57.03%的f分,这是此类任务中报道的最高结果。考虑到参与的水平和团队的结果,我们发现我们的任务在文本挖掘研究社区的参与上是成功的,产生了一个大型的带注释的语料库,并改善了自动疾病识别和CDR提取的结果。
英文摘要
Mining useful knowledge from the biomedical literature holds potentials for helping literature searching, automating biological data curation and many other scientific tasks. Hence, it is important to be able to recognize various types of biological entities in free text, such as gene/proteins, disease/conditions, and drug/chemicals, etc. Indeed, our previous PubMed log analysis revealed that people search certain biomedical concepts more often than others and that there exist strong associations between different concepts. For example, a disease name often co-occurs with gene/proteins and drug names. To assess the state of the art in biomedical entity recognition and relation extraction, we organized a science competition at BioCreative V, an international challenge event for evaluating advances in text mining research for biology. Specifically, we designed two challenge tasks: disease named entity recognition (DNER) and chemical-induced disease (CID) relation extraction. To assist system development and assessment, we created a large annotated text corpus that consisted of human annotations of chemicals, diseases and their interactions from 1500 PubMed articles. 34 teams worldwide participated in the CDR task: 16 (DNER) and 18 (CID). The best systems achieved an F-score of 86.46% for the DNER task--a result that approaches the human inter-annotator agreement (0.8875)--and an F-score of 57.03% for the CID task, the highest results ever reported for such tasks. Given the level of participation and team results, we found our task to be successful in engaging the text-mining research community, producing a large annotated corpus and improving the results of automatic disease recognition and CDR extraction. In addition to organizing the BioCreative task, we continued our own development of biomedical named entity taggers in 2015-2016. First and foremost, we created a general toolkit called TaggerOne: the first machine learning model for joint named entity recognition and normalization. TaggerOne is an all-purpose tagger (i.e. not specific to any entity type), requiring only annotated training data and a corresponding lexicon, and has been optimized for high throughput. We validated TaggerOne with multiple gold-standard corpora containing both mention- and concept-level annotations. Its results compare favorably to the previous state of the art, notwithstanding the greater flexibility of the model. TaggerOne is implemented in Java and its source code has been made publicly available to the research community. However, large-scale use of open-source tools sometimes requires a significant investment in infrastructure and maintenance time. These investments not only impair the continued adoption of text mining tools, but also reduce the ability of individual researchers to explore applying text mining to problems in their research area. In contrast, Web services provide on-demand access to software tools through the Internet using straightforward interfaces and data formats. Providing text mining tools as web services therefore reduces the bar to use for biocurators and bioinformatics researchers not working specifically in text mining, allowing free exploration and the ability to focus on results rather than methodology. Therefore, in 2015 we developed NCBI text-mining web services, an online version of our text mining tool suite for biomedical concept recognition and information extraction. Our service incorporates multiple state of the art tools for identifying critical entity types: DNorm (for diseases), GNormPlus (genes and proteins), SR4GN (species), tmChem (chemicals and drugs), and tmVar (variants). Our web service has already processed over 60 million requests since its inception from researchers in 46 countries, supporting research projects in biocuration, crowdsourcing and translational bioinformatics. We anticipate that providing text mining tools as web services will greatly expand their utility to the biomedical research community. Finally, as mentioned earlier, one promising application area for text mining research is to assist manual literature curation, a highly time-consuming and labor-intensive process. In this regard, we continued to improve our previous curation-assisting tool PubTator and to collaborate with domain experts: human database curators in this case. With these efforts, our PubTator system is continuously being used in the production curation pipeline of two external databases on a daily basis: 1. HuGE Navigator a CDCs knowledgebase of human genome epidemiology 2. SwissProt an annotated database of protein sequence and functional information
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Query Log Analysis for Improving User Access to NCBI Web Services
  • 批准号:
    9564626
  • 项目类别:
  • 资助金额:
    $160.63万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Machine Learning and Natural Language Processing for Biomedical Applications
  • 批准号:
    10927050
  • 项目类别:
  • 资助金额:
    $387.34万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
  • 批准号:
    10007525
  • 项目类别:
  • 资助金额:
    $190.14万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
  • 批准号:
    8149607
  • 项目类别:
  • 资助金额:
    $39.17万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
海外基金