Classification of Full Text Biomedical Documents: Sections Importance Assessment

Classification of Full Text Biomedical Documents: Sections Importance Assessment
复制标题

DOI:
10.3390/app11062674
复制
发表时间:
2021-03-01
影响因子:
2.7
通讯作者:
Lorenzo Iglesias, Eva
Lorenzo Iglesias, Eva
中科院分区:
综合性期刊4区
文献类型:
--
作者:
Oliveira Goncalves, Carlos Adriano;Camacho, Rui;Lorenzo Iglesias, Eva

文献摘要

被引文献

相似文献

网络上文档的指数级增长使得研究人员很难了解科学界正在进行的相关工作。因此,有效地检索信息的任务已成为一个重要的研究课题。本研究的目的是测试如何的效率的文本分类的变化,如果不同的权重先前分配给组成的文件的部分。该建议考虑了术语在文档中的位置(部分),每个部分都有一个权重,可以根据语料库进行修改。为了进行这项研究,我们创建了包含完整文档的OHSUMED语料库的扩展版本。通过使用WEKA,我们比较了仅使用摘要和全文,以及使用章节加权组合,以评估其在使用SMO(序列最小优化),WEKA支持向量机(SVM)算法实现的科学文章分类过程中的意义。实验结果表明,所提出的预处理技术和特征选择的组合取得了可喜的成果,全文科学文献分类的任务。我们也有证据得出结论,从某些部分的文本丰富的数据集实现更好的结果比只使用标题和摘要。
The exponential growth of documents in the web makes it very hard for researchers to be aware of the relevant work being done within the scientific community. The task of efficiently retrieving information has therefore become an important research topic. The objective of this study is to test how the efficiency of the text classification changes if different weights are previously assigned to the sections that compose the documents. The proposal takes into account the place (section) where terms are located in the document, and each section has a weight that can be modified depending on the corpus. To carry out the study, an extended version of the OHSUMED corpus with full documents have been created. Through the use of WEKA, we compared the use of abstracts only with that of full texts, as well as the use of section weighing combinations to assess their significance in the scientific article classification process using the SMO (Sequential Minimal Optimization), the WEKA Support Vector Machine (SVM) algorithm implementation. The experimental results show that the proposed combinations of the preprocessing techniques and feature selection achieve promising results for the task of full text scientific document classification. We also have evidence to conclude that enriched datasets with text from certain sections achieve better results than using only titles and abstracts.