课题基金 / 基金详情

项目摘要

项目成果

Zhiyong Lu的其他基金

相似基金

相关文献

中文摘要
翻译
从生物医学文献中挖掘有用的知识有助于文献检索,自动化生物数据管理和许多其他科学任务。因此,重要的是能够识别自由文本中的各种类型的生物实体,例如基因/蛋白质,疾病/病症和药物/化学品等。事实上,我们以前的PubMed日志分析显示,人们搜索某些生物医学概念的频率高于其他概念,并且不同概念之间存在很强的关联。例如,疾病名称通常与基因/蛋白质和药物名称共同出现。我们最近的研究介绍了一种名为DNorm的最先进的系统,用于基于成对学习排名的疾病标准化。在2013-2014年,我们进行了一项关于DNorm应用于临床叙述与生物医学出版物时的不同性能的调查。我们使用闭包属性来比较临床叙述文本与生物医学出版物中词汇的丰富程度。我们发现,虽然临床叙述和生物医学出版物之间的总体词汇量相似,但临床叙述使用比出版物更丰富的术语来描述疾病,我们认为这是临床叙述性能降低的主要原因之一。因此,我们引入了几个可推广到其他临床NLP任务的词汇增强,这些任务提高了DNorm处理这种变化的能力。DNorm的临床版本(DNorm-C)现在与我们的其他开源工具一起向研究社区开放,沿着。 生物医学命名实体识别(NER)和标准化中的一个常见挑战是复合命名实体的识别和解析,其中单个跨度指代多于一个概念(例如,BRCA 1/2)。以前的NER和规范化研究要么忽略复合提及,使用简单的特设规则,或只处理协调省略,使得处理多类型复合提及非常需要一个强大的方法。在2014-2015年,我们提出了一种混合方法,将机器学习模型与模式识别策略相结合,以识别每个复合提及的各个组成部分。我们的方法,我们命名为SimConcept,是第一个系统地处理多种类型的复合提及。该技术在识别和解决三个关键生物实体的复合提及方面取得了很高的性能:基因(90.42%的F-测度),疾病(86.47%的F-测度)和化学物质(86.05%的F-测度)。此外,我们的结果表明,使用我们的SimConcept方法可以随后提高基因和疾病概念识别和规范化的性能。 如前所述,文本挖掘研究的一个很有前途的应用领域是帮助手动文献管理,这是一个非常耗时和劳动密集型的过程。在这方面,我们继续改进我们以前的策展辅助工具PubTator,并与领域专家合作:在这种情况下是人类数据库策展者。通过这些努力,我们的PubTator系统现在每天都用于两个外部数据库的生产策展管道: 1. HuGE Navigator:CDC人类基因组流行病学知识库 2. SwissProt:蛋白质序列和功能信息的注释数据库。 在2014-2015年,我们还研究了利用众包分别协助基因突变治疗和药物适应症编目的可行性,因为专家注释的成本很高。在这两项研究中,我们首先将复杂的专家注释任务转化为适合普通工人的人类智力任务(HIT)。例如,我们没有要求人们从自由文本(例如冗长的段落)中找到药物适应症,而是简化了任务,使得每个HIT只涉及一个工人,在给定的药物标签句子的上下文中,对突出显示的疾病是否是适应症进行二元判断。然后,我们通过Amazon Mechanical Turk(MTurk)的技术环境从未知的工作人员网络中招募注释者。众人的评价,再汇总,成为最终的答案。为了进行评估,我们评估了我们提出的方法以具有时间效率和成本效益的方式实现高质量注释的能力。与专家注释相比,我们发现,我们的众包方法不仅节省了大量的成本和时间,而且还导致与领域专家的准确性相当。因此,我们得出结论,我们的基于众包的方法提供了一个易于扩展和具有成本效益的模型,手动策展。
英文摘要
Mining useful knowledge from the biomedical literature holds potentials for helping literature searching, automating biological data curation and many other scientific tasks. Hence, it is important to be able to recognize various types of biological entities in free text, such as gene/proteins, disease/conditions, and drug/chemicals, etc. Indeed, our previous PubMed log analysis revealed that people search certain biomedical concepts more often than others and that there exist strong associations between different concepts. For example, a disease name often co-occurs with gene/proteins and drug names. Our recent research introduced a state-of-the-art system called DNorm for disease normalization based on pairwise learning to rank. In 2013-2014, we performed an investigation with regard to the different performance of DNorm when applied to clinical narratives vs. biomedical publications. We used closure properties to compare the richness of the vocabulary in clinical narrative text to biomedical publications. We found that while the size of the overall vocabulary is similar between clinical narrative and biomedical publications, clinical narrative uses a richer terminology to describe disorders than publications, which we believe to be one of the primary causes of reduced performance in clinical narrative. Accordingly, we introduced several lexical enhancements generalizable to other clinical NLP tasks that improved the ability of DNorm to handle this variation. The clinical version of DNorm (DNorm-C) is now made openly available to the research community, along with our other open source tools. One common challenge in biomedical named entity recognition (NER) and normalization is the identification and resolution of composite named entities, where a single span refers to more than one concept (e.g., BRCA1/2). Previous NER and normalization studies have either ignored composite mentions, used simple ad hoc rules, or only handled coordination ellipsis, making a robust approach for handling multitype composite mentions greatly needed. In 2014-2015, we proposed a hybrid method integrating a machine-learning model with a pattern identification strategy to identify the individual components of each composite mention. Our method, which we have named SimConcept, is the first to systematically handle many types of composite mentions. The technique achieves high performance in identifying and resolving composite mentions for three key biological entities: genes (90.42% in F-measure), diseases (86.47% in F-measure), and chemicals (86.05% in F-measure). Furthermore, our results show that using our SimConcept method can subsequently improve the performance of gene and disease concept recognition and normalization. As mentioned earlier, one promising application area for text mining research is to assist manual literature curation, a highly time-consuming and labor-intensive process. In this regard, we continued to improve our previous curation-assisting tool PubTator and to collaborate with domain experts: human database curators in this case. With these efforts, our PubTator system is now being used in the production curation pipeline of two external databases on a daily basis: 1. HuGE Navigator: a CDC knowledgebase of human genome epidemiology 2. SwissProt: an annotated database of protein sequence and functional information. In 2014-2015, we also investigated the feasibility of using crowdsourcing for respectively assisting gene-mutation curation and drug-indication cataloging, given the high cost of expert annotation. In both studies, we first translated the complex expert-annotation task into human intelligence tasks (HITs) suitable for the average workers. For instance, instead of asking people to find drug indications from free text (e.g. lengthy paragraphs), we simplified the task such that each HIT only involved a worker making a binary judgment of whether a highlighted disease, in context of a given drug label sentence, is an indication. Then we recruited annotators from an unknown network of workers through the technical environment of Amazon Mechanical Turk (MTurk). Judgments from the crowds were then aggregated to become the final answer. For evaluation, we assessed the ability of our proposed method to achieve high-quality annotations in a time-efficient and cost-effective manner. In comparison with the expert annotations, we find that our crowdsourcing approach not only results in significant cost and time saving, but also leads to accuracy comparable to that of domain experts. Therefore, we conclude that our crowdsourcing-based approach provides a readily scalable and cost-effective model to manual curation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Named Entity Recognition and Relationship Extraction in Biomedicine
  • 批准号:
    9362446
  • 项目类别:
  • 资助金额:
    $140.39万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
  • 批准号:
    9564626
  • 项目类别:
  • 资助金额:
    $160.63万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Machine Learning and Natural Language Processing for Biomedical Applications
  • 批准号:
    10927050
  • 项目类别:
  • 资助金额:
    $387.34万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
  • 批准号:
    10007525
  • 项目类别:
  • 资助金额:
    $190.14万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
  • 批准号:
    2021JJ40433
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2021
  • 负责人:
    孙磊
  • 依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
  • 批准号:
    32001603
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    段真珍
  • 依托单位:
AREA国际经济模型的移植.改进和应用
  • 批准号:
    18870435
  • 项目类别:
    面上项目
  • 资助金额:
    2.0万元
  • 批准年份:
    1988
  • 负责人:
    史树中
  • 依托单位: