A Novel Multi-View Ensemble Learning Architecture to Improve the Structured Text Classification

A Novel Multi-View Ensemble Learning Architecture to Improve the Structured Text Classification
复制标题

DOI:
10.3390/info13060283
复制
发表时间:
2022-06-01
期刊:
影响因子:
3.1
通讯作者:
Diz, Lourdes Borrajo
Diz, Lourdes Borrajo
中科院分区:
其他
文献类型:
--
作者:
Goncalves, Carlos Adriano;Vieira, Adrian Seara;Diz, Lourdes Borrajo

文献摘要

被引文献

相似文献

多视图集成学习利用数据视图的信息。为了测试其对全文分类的效率,已经实现了一种技术,其中视图对应于文档部分。对于分类和预测,我们使用基于不同学习算法提供数据互补解释的思想的堆叠泛化。本研究采用支持向量机算法作为基线,C4.5实现作为元学习者,实现了堆栈方法。使用OHSUMED生物医学全文文档创建视图。实验结果表明,将多视图技术应用于全文文本,显著提高了文本分类的效率,为生物医学文本挖掘研究做出了重要贡献。我们也有证据得出结论,从某些部分的文本丰富的数据集比只使用标题和摘要更好。
Multi-view ensemble learning exploits the information of data views. To test its efficiency for full text classification, a technique has been implemented where the views correspond to the document sections. For classification and prediction, we use a stacking generalization based on the idea that different learning algorithms provide complementary explanations of the data. The present study implements the stacking approach using support vector machine algorithms as the baseline and a C4.5 implementation as the meta-learner. Views are created with OHSUMED biomedical full text documents. Experimental results lead to the sustained conclusion that the application of multi-view techniques to full texts significantly improves the task of text classification, providing a significant contribution for the biomedical text mining research. We also have evidence to conclude that enriched datasets with text from certain sections are better than using only titles and abstracts.