Improving imbalanced scientific text classification using sampling strategies and dictionaries

Improving imbalanced scientific text classification using sampling strategies and dictionaries
复制标题

DOI:
10.2390/biecoll-jib-2011-176
复制
发表时间:
2011-12-01
影响因子:
1.9
通讯作者:
Redondo Marey, C. M.
Redondo Marey, C. M.
中科院分区:
其他
文献类型:
--
作者:
Borrajo, L.;Romero, R.;Redondo Marey, C. M.

文献摘要

被引文献

相似文献

许多实际应用程序都存在不平衡的类分布问题,与其他类相比,其中一个类的用例数量非常少。受影响的系统之一是与科学文件的恢复和分类有关的系统。过采样和次采样等采样策略是解决类不平衡问题的常用方法。在这项工作中,我们研究了它们在PubMed科学数据库上搜索时对三种分类器(Knn, SVM和Naive-Bayes)的影响。本文的另一个目的是研究词典在生物医学文本分类中的应用。实验使用了三个不同的字典(BioCreative、NLPBA和UniProt数据库的一个名为Protein的特别子集),使用了上述分类器和采样策略。使用NLPBA和Protein字典以及使用Subsampling平衡技术的SVM分类器获得了最好的结果。这些结果与其他作者使用TREC Genomics 2005公共语料库获得的结果进行了比较。
Many real applications have the imbalanced class distribution problem, where one of the classes is represented by a very small number of cases compared to the other classes. One of the systems affected are those related to the recovery and classification of scientific documentation.Sampling strategies such as Oversampling and Subsampling are popular in tackling the problem of class imbalance. In this work, we study their effects on three types of classifiers (Knn, SVM and Naive-Bayes) when they are applied to search on the PubMed scientific database.Another purpose of this paper is to study the use of dictionaries in the classification of biomedical texts. Experiments are conducted with three different dictionaries (BioCreative, NLPBA, and an ad-hoc subset of the UniProt database named Protein) using the mentioned classifiers and sampling strategies.Best results were obtained with NLPBA and Protein dictionaries and the SVM classifier using the Subsampling balancing technique. These results were compared with those obtained by other authors using the TREC Genomics 2005 public corpus.