A Framework for Semisupervised Feature Generation and Its Applications in Biomedical Literature Mining

A Framework for Semisupervised Feature Generation and Its Applications in Biomedical Literature Mining
复制标题

DOI:
10.1109/tcbb.2010.99
复制
发表时间:
2011-03-01
影响因子:
4.5
通讯作者:
Yang, Zhihao
Yang, Zhihao
中科院分区:
工程技术3区
文献类型:
--
作者:
Li, Yanpeng;Hu, Xiaohua;Yang, Zhihao

文献摘要

被引文献

相似文献

特征表示对于机器学习和文本挖掘至关重要。在本文中,我们提出了一种特征耦合泛化(FCG)框架,用于从未标记的数据生成新特征。它从原始特征集中选择两种特殊类型的特征,即示例区分特征(EDF)和类区分特征(CDF),然后根据 EDF 与未标记数据中 CDF 的耦合程度将 EDF 泛化为更高级别的特征。其优点是:标记数据中极度稀疏的EDF可以通过与未标记数据中的CDF共现来丰富,从而可以大大提高这些低频特征的性能,并可以合并来自未标记的新信息。我们将该方法应用于生物医学文献挖掘中的三个任务:基因命名实体识别(NER)、蛋白质-蛋白质相互作用提取(PPIE)和用于基因本体(GO)注释的文本分类(TC)。新特征是从超过 20 GB 未标记的 PubMed 摘要中生成的。 BioCreative 2、AIMED 语料库和 TREC 2005 Genomics Track 上的实验结果表明:1)FCG 可以很好地利用监督学习忽略的稀疏特征。 2)它在树任务中将监督基线的性能分别提高了 7.8%、5.0% 和 5.8%。 3) 我们的方法在三个基准数据集上实现了 89.1、64.5 F 分数和 60.1 标准化效用。
Feature representation is essential to machine learning and text mining. In this paper, we present a feature coupling generalization (FCG) framework for generating new features from unlabeled data. It selects two special types of features, i.e., example-distinguishing features (EDFs) and class-distinguishing features (CDFs) from original feature set, and then generalizes EDFs into higher-level features based on their coupling degrees with CDFs in unlabeled data. The advantage is: EDFs with extreme sparsity in labeled data can be enriched by their co-occurrences with CDFs in unlabeled data so that the performance of these low-frequency features can be greatly boosted and new information from unlabeled can be incorporated. We apply this approach to three tasks in biomedical literature mining: gene named entity recognition (NER), protein-protein interaction extraction (PPIE), and text classification (TC) for gene ontology (GO) annotation. New features are generated from over 20 GB unlabeled PubMed abstracts. The experimental results on BioCreative 2, AIMED corpus, and TREC 2005 Genomics Track show that 1) FCG can utilize well the sparse features ignored by supervised learning. 2) It improves the performance of supervised baselines by 7.8 percent, 5.0 percent, and 5.8 percent, respectively, in the tree tasks. 3) Our methods achieve 89.1, 64.5 F-score, and 60.1 normalized utility on the three benchmark data sets.