Discriminative and informative features for biomolecular text mining with ensemble feature selection.

Discriminative and informative features for biomolecular text mining with ensemble feature selection.
复制标题

DOI:
10.1093/bioinformatics/btq381
复制
发表时间:
2010-09-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Van de Peer Y
Van de Peer Y
中科院分区:
其他
文献类型:
--
作者:
Van Landeghem S;Abeel T;Saeys Y;Van de Peer Y

文献摘要

参考文献

被引文献

相似文献

动机:在生物分子文本挖掘领域,机器学习系统的黑箱行为目前限制了对预测真实性质的理解。然而,特征选择(FS)能够在任何监督学习环境中识别最相关的特征,从而提供对分类算法的特定属性的洞察。这使我们能够构建更准确的分类器,同时在黑盒行为和必须解释结果的最终用户之间架起一座桥梁。结果:我们的FS方法成功地丢弃了很大一部分机器生成的特征,提高了最新文本挖掘算法的分类性能。此外,我们还说明了FS如何应用于从文本中提取生物分子事件的框架的预测中获得理解。我们列举了大量具有高度区别性的特征的例子,这些特征要么模拟生物现实,要么模拟常见的语言结构。最后,我们讨论了来自FS分析的一些见解,这些见解将为显著改进当前的文本挖掘工具提供机会。可用性:FS算法和分类器在JAVA-ML(http://java-ml.sf.net).)中可用这些数据集可从BioNLP‘09共享任务网站(http://www-tsujii.is.s.u-tokyo.ac.jp/GENIA/SharedTask/).公开获得联系人:yves.vandepeer@psb.ugent.be
Motivation: In the field of biomolecular text mining, black box behavior of machine learning systems currently limits understanding of the true nature of the predictions. However, feature selection (FS) is capable of identifying the most relevant features in any supervised learning setting, providing insight into the specific properties of the classification algorithm. This allows us to build more accurate classifiers while at the same time bridging the gap between the black box behavior and the end-user who has to interpret the results. Results: We show that our FS methodology successfully discards a large fraction of machine-generated features, improving classification performance of state-of-the-art text mining algorithms. Furthermore, we illustrate how FS can be applied to gain understanding in the predictions of a framework for biomolecular event extraction from text. We include numerous examples of highly discriminative features that model either biological reality or common linguistic constructs. Finally, we discuss a number of insights from our FS analyses that will provide the opportunity to considerably improve upon current text mining tools. Availability: The FS algorithms and classifiers are available in Java-ML (http://java-ml.sf.net). The datasets are publicly available from the BioNLP'09 Shared Task web site (http://www-tsujii.is.s.u-tokyo.ac.jp/GENIA/SharedTask/). Contact: yves.vandepeer@psb.ugent.be
评估生物学的文本挖掘系统:第二次生物综合社区挑战的概述。
DOI: 10.1186/gb-2008-9-s2-s1
发表时间: 2008
期刊: Genome biology
影响因子: 12.3
作者:
Krallinger M;Morgan A;Smith L;Leitner F;Tanabe L;Wilbur J;Hirschman L;Valencia A
通讯作者: Valencia A
通过评估跨核心学习,全路径图核用于蛋白质 - 蛋白质相互作用提取。
DOI: 10.1186/1471-2105-9-s11-s2
发表时间: 2008-11-19
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Airola, Antti;Pyysalo, Sampo;Bjoerne, Jari;Pahikkala, Tapio;Ginter, Filip;Salakoski, Tapio
通讯作者: Salakoski, Tapio
DOI: 10.1093/bioinformatics/btp191
发表时间: 2009-06-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Abeel T;Van de Peer Y;Saeys Y
通讯作者: Saeys Y
DOI: 10.1108/eb046814
发表时间: 2006-01-01
影响因子: --
作者:
Porter, M. F.
通讯作者: Porter, M. F.
DOI: 10.1186/1756-0381-1-8
发表时间: 2008-09-19
期刊: BioData mining
影响因子: 4.5
作者:
Reverter A;Ingham A;Dalrymple BP
通讯作者: Dalrymple BP