Incorporating rich background knowledge for gene named entity classification and recognition.

Incorporating rich background knowledge for gene named entity classification and recognition.
复制标题

结合丰富的背景知识进行基因命名实体分类和识别

DOI:
10.1186/1471-2105-10-223
复制
发表时间:
2009-07-17
期刊:
影响因子:
3
通讯作者:
Yang Z
Yang Z
中科院分区:
生物学4区
文献类型:
--
作者:
Li Y;Lin H;Yang Z

文献摘要

参考文献

被引文献

相似文献

基因命名实体的分类与识别是生物医学文献文本挖掘的关键步骤。基于机器学习的方法已经在这一领域获得了巨大的成功。在大多数最先进的系统中,精心设计的词汇特征,如单词、n-gram和形态学模式,发挥了核心作用。然而,这种类型的特征往往会导致特征空间的极度稀疏。因此,由于缺乏信息,训练数据中的词汇外(OOV)术语不能很好地建模。我们提出了一种通用的基因命名实体表示框架,称为特征耦合泛化(FCG)。其基本思想是利用大量未标记数据中具有高度指示性的特征的词频和共现信息来生成更高层次的特征。我们在一个命名实体分类任务中检验了它的性能,该任务旨在删除来自在线资源的大型字典中的非基因条目。结果表明,FCG生成的新特征比词汇特征的f值高5.97分,比OOV词的f值高10.85分。此外,在这个框架中,每个扩展都会产生显著的改进,并且稀疏的词法特征可以转换为更低维度和更有信息的表示。基于改进字典的前向最大匹配方法在BioCreative 2 GM测试集上获得了86.2的f分。然后,我们将词典与基于条件随机场(CRF)的基因提及标注器相结合,获得了89.05的f分,这使得基于条件随机场(CRF)的标注器的性能提高了4.46,对识别系统的效率影响很小。NER系统的演示可在。
Gene named entity classification and recognition are crucial preliminary steps of text mining in biomedical literature. Machine learning based methods have been used in this area with great success. In most state-of-the-art systems, elaborately designed lexical features, such as words, n-grams, and morphology patterns, have played a central part. However, this type of feature tends to cause extreme sparseness in feature space. As a result, out-of-vocabulary (OOV) terms in the training data are not modeled well due to lack of information. We propose a general framework for gene named entity representation, called feature coupling generalization (FCG). The basic idea is to generate higher level features using term frequency and co-occurrence information of highly indicative features in huge amount of unlabeled data. We examine its performance in a named entity classification task, which is designed to remove non-gene entries in a large dictionary derived from online resources. The results show that new features generated by FCG outperform lexical features by 5.97 F-score and 10.85 for OOV terms. Also in this framework each extension yields significant improvements and the sparse lexical features can be transformed into both a lower dimensional and more informative representation. A forward maximum match method based on the refined dictionary produces an F-score of 86.2 on BioCreative 2 GM test set. Then we combined the dictionary with a conditional random field (CRF) based gene mention tagger, achieving an F-score of 89.05, which improves the performance of the CRF-based tagger by 4.46 with little impact on the efficiency of the recognition system. A demo of the NER system is available at .
DOI: 10.1093/bioinformatics/bti475
发表时间: 2005-07-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Settles, B
通讯作者: Settles, B
DOI: 10.1186/1471-2105-6-s1-s2
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Yeh A;Morgan A;Colosimo M;Hirschman L
通讯作者: Hirschman L
DOI: 10.1142/s0219720004000399
发表时间: 2004-01-01
影响因子: 1
作者:
Tanabe, Lorraine;Wilbur, W John
通讯作者: Wilbur, W John
DOI: 10.1093/bioinformatics/btn299
发表时间: 2008-08-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hakenberg, Joerg;Plake, Conrad;Gonzalez, Graciela
通讯作者: Gonzalez, Graciela
DOI: 10.1016/j.artint.2005.03.001
发表时间: 2005-06-01
影响因子: 14.4
作者:
Etzioni, O;Cafarella, M;Yates, A
通讯作者: Yates, A