Generation of a large gene/protein lexicon by morphological pattern analysis.

Generation of a large gene/protein lexicon by morphological pattern analysis.
复制标题

DOI:
10.1142/s0219720004000399
复制
发表时间:
2004-01-01
影响因子:
1
通讯作者:
Wilbur, W John
Wilbur, W John
中科院分区:
生物学4区
文献类型:
--
作者:
Tanabe, Lorraine;Wilbur, W John

文献摘要

被引文献

相似文献

自然语言文本中基因/蛋白质名称的识别是命名实体识别中的一个重要问题。在以前的工作中,我们处理了MEDLINE文件,获得了超过200万个名称的集合,我们估计其中可能有三分之二是有效的基因/蛋白质名称。我们的问题一直是如何纯化这一套,以获得高质量的基因/蛋白质名称的子集。在这里,我们描述了一种方法,这是基于生成的某些类别的名称,其特征在于共同的形态特征。在每个类中,应用归纳逻辑编程(ILP)来学习那些基因/蛋白质名称的特征。以这种方式学习的标准然后应用于我们的大集合的名称。我们生成了193类名称,ILP导致定义1,240,462个名称的选择子集的标准。一个简单的假阳性过滤器被应用于删除8%的这个集合留下1,145,913个名字。对来自该基因/蛋白质名称词典的随机样本的检查表明,它由82%(+/-3%)完整和准确的基因/蛋白质名称,12%与基因/蛋白质相关的名称(太通用,有效名称加上附加文本,有效名称的一部分等),6%的名称与基因/蛋白质无关。该词典可在ftp.ncbi.nlm.nih.gov/pub/tanabe/Gene.Lexicon上免费获得。
The identification of gene/protein names in natural language text is an important problem in named entity recognition. In previous work we have processed MEDLINE documents to obtain a collection of over two million names of which we estimate that perhaps two thirds are valid gene/protein names. Our problem has been how to purify this set to obtain a high quality subset of gene/protein names. Here we describe an approach which is based on the generation of certain classes of names that are characterized by common morphological features. Within each class inductive logic programming (ILP) is applied to learn the characteristics of those names that are gene/protein names. The criteria learned in this manner are then applied to our large set of names. We generated 193 classes of names and ILP led to criteria defining a select subset of 1,240,462 names. A simple false positive filter was applied to remove 8% of this set leaving 1,145,913 names. Examination of a random sample from this gene/protein name lexicon suggests it is composed of 82% (+/-3%) complete and accurate gene/protein names, 12% names related to genes/proteins (too generic, a valid name plus additional text, part of a valid name, etc.), and 6% names unrelated to genes/proteins. The lexicon is freely available at ftp.ncbi.nlm.nih.gov/pub/tanabe/Gene.Lexicon.