课题基金 / 基金详情

项目摘要

项目成果

Zhiyong Lu的其他基金

相似基金

相关文献

中文摘要
翻译
从生物医学文献中挖掘有用的知识具有帮助文献搜索、自动化生物数据管理和许多其他科学任务的潜力。因此,我们专注于识别自由文本中的各种类型的生物实体,如基因/蛋白质、疾病/疾病和药物/化学品等,以及它们之间的关系。 同义词对高质量的相关性搜索提出了另一个挑战。这对于普通单词来说是一个问题,但对于可以用多种不同方式命名的实体来说,这就更难了。LitVar针对基因变异解决了这个问题。例如,搜索A146T、C.436G>A或rs121913527中的一个也会找到其他两个的实例。目标是将此功能扩展到其他实体类型。 我们在BioCreative VI上参加了CHEMPROT跟踪,该跟踪旨在评估在运行文本(PubMed摘要)中自动提取化学蛋白质关系的技术水平。我们提出了一个由三个系统组成的集成,包括支持向量机、卷积神经网络和递归神经网络。他们的输出使用多数投票或堆叠的方式进行组合,以进行最终预测。我们的系统在挑战期间获得了0.7266的准确率和0.5735的召回率,F-分数为0.6410,在所有团队提交的挑战中取得了最高的表现。 除了用有监督的机器学习方法处理关系提取任务外,我们还提出了一种新的对抗性学习算法,用于目标领域没有标签数据的无监督领域自适应任务。我们证明了领域不变特征可以在最新的神经网络中学习,这样为一种关系类型(蛋白质)训练的分类器可以重新用于其他关系类型(毒品)。与以往基于卷积和递归神经网络的关系分类方法相比,在没有领域自适应的情况下,我们的F1-Score提高了30%。为了在没有预先存在的训练数据的情况下进一步帮助NLP任务,我们开发了ezTag,这是一个基于Web的标注工具,允许用户执行标注并在循环中提供训练数据。EzTag支持PubMed中的摘要和PubMed Central中的全文文章。 阴性和不确定的医学发现在放射学报告中经常出现,但将它们与阳性发现区分开来仍然是信息提取的挑战。在这里,我们提出了一种新的算法,NegBio,来检测放射学报告中的阴性和不确定的发现。与以前基于规则的方法不同,NegBio利用通用依赖模式来识别指示否定或不确定性的触发器的范围。我们在四个数据集上对NegBio进行了评估,包括两个放射学报告的公共基准语料库,我们为这项工作注释的一个新的放射学语料库,以及一个公共的一般临床文本语料库。对这些数据集的评估表明,NegBio在检测负面和不确定的发现方面非常准确,与当前的技术水平相比是有利的。 文本挖掘研究的一个很有前途的应用领域是辅助手动文献整理,这是一个非常耗时和劳动密集型的过程。在这方面,我们通过与UniProtKB/Swiss-Prot和NHGRI-EBI Gwas Catalog的数据库馆长合作,将自动深度学习技术应用于其基因组变异的文献分类过程。两个人工管理团队都证实,我们的方法在不影响召回率的情况下,实现了比他们之前的基于查询的分类方法更高的精度。实验结果表明,该方法具有较高的效率,可以取代传统的基于查询的人工分类数据库分类方法。我们的方法可以让人类策展人有更多的时间专注于更具挑战性的任务,例如实际的策展以及发现新的论文/实验技术以考虑纳入。 深度学习是机器学习算法的一类,在我们最近的几项研究中显示了令人印象深刻的结果,如上在18财年所示。除了在自然语言处理中的应用外,我们还看到它在医学图像分析中的成功,例如处理胸部X光图像和彩色眼底照片。
英文摘要
Mining useful knowledge from the biomedical literature holds potentials for helping literature searching, automating biological data curation and many other scientific tasks. We have therefore focused on recognizing various types of biological entities in free text, such as gene/proteins, disease/conditions, and drug/chemicals, etc, and their relationships. Synonyms pose another challenge for high quality relevance searches. This is a problem for ordinary words, but it is even more of a difficult for entities that can be named in a number of different ways. LitVar address this problem for genetic variants. For example, searching for one of A146T, c.436G>A, or rs121913527 also finds instances of the other two. The goal is to extend this ability to other entity types. We participated in The CHEMPROT track at BioCreative VI, which aims to assess the state of the art in automatically extracting the chemicalprotein relations in running text (PubMed abstracts). We proposed an ensemble of three systems, including a support vector machine, a convolutional neural network, and a recurrent neural network. Their output is combined using majority voting or stacking for final predictions. Our system obtained 0.7266 in precision and 0.5735 in recall for an F-score of 0.6410 during the challenge, achieving the highest performance among all team submissions during the challenge. In addition to tackling relation extraction tasks with supervised machine-learning methods, we proposed a novel adversarial learning algorithm for unsupervised domain adaptation tasks where no labeled data are available in the target domain. We show domain invariant features can be learned in the latest neural networks such that classifiers trained for one relation type (proteinprotein) can be re-purposed to others (drugdrug). Compared to prior convolutional and recurrent NN-based relation classification methods without domain adaptation, we achieve improvements as high as 30% in F1-score. To further assist NLP tasks without pre-existing training data, we developed ezTag, a web-based annotation tool that allows users to perform annotation and provide training data with humans in the loop. ezTag supports both abstracts in PubMed and full-text articles in PubMed Central. Negative and uncertain medical findings are frequent in radiology reports, but discriminating them from positive findings remains challenging for information extraction. Here, we propose a new algorithm, NegBio, to detect negative and uncertain findings in radiology reports. Unlike previous rule-based methods, NegBio utilizes patterns on universal dependencies to identify the scope of triggers that are indicative of negation or uncertainty. We evaluated NegBio on four datasets, including two public benchmarking corpora of radiology reports, a new radiology corpus that we annotated for this work, and a public corpus of general clinical texts. Evaluation on these datasets demonstrates that NegBio is highly accurate for detecting negative and uncertain findings and compares favorably to the current state of the art. One promising application area for text mining research is to assist manual literature curation, a highly time-consuming and labor-intensive process. In this regard, we applied automated deep learning techniques to the literature triage process of UniProtKB/Swiss-Prot and the NHGRI-EBI GWAS Catalog for genomic variation by collaborating with their database curators. Both the manual curation teams confirmed that our method achieved higher precision than their previous query-based triage methods without compromising recall. Both results show that our method is more efficient and can replace the traditional query-based triage methods of manually curated databases. Our method can give human curators more time to focus on more challenging tasks such as actual curation as well as the discovery of novel papers/experimental techniques to consider for inclusion. Deep learning, a class of machine learning algorithms, has showed impressive results in several of our recent studies as shown above in FY18. In addition to its applications in natural language processing, we have also seen its success in our medical image analysis such as processing chest X-ray images and colors fundus photographs.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Named Entity Recognition and Relationship Extraction in Biomedicine
  • 批准号:
    9362446
  • 项目类别:
  • 资助金额:
    $140.39万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
  • 批准号:
    9564626
  • 项目类别:
  • 资助金额:
    $160.63万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Machine Learning and Natural Language Processing for Biomedical Applications
  • 批准号:
    10927050
  • 项目类别:
  • 资助金额:
    $387.34万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
  • 批准号:
    10007525
  • 项目类别:
  • 资助金额:
    $190.14万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
海外基金