Active learning for clinical text classification: is it better than random sampling?

Active learning for clinical text classification: is it better than random sampling?
复制标题

DOI:
10.1136/amiajnl-2011-000648
复制
发表时间:
2012-09-01
影响因子:
6.4
通讯作者:
Wiechmann, Eduardo P.
Wiechmann, Eduardo P.
中科院分区:
管理学2区
文献类型:
--
作者:
Figueroa, Rosa L.;Zeng-Treitler, Qing;Wiechmann, Eduardo P.

文献摘要

被引文献

相似文献

目的探讨主动学习算法在医学文本分类中的应用,以减少对大样本集的需求。设计三种现有的主动学习算法(基于距离(distance based,DIST)、基于多样性(diversity-based,DIV)和两者结合(combination of both,CMB))分别用于5个数据集的文本分类。这些算法的性能进行了比较,被动学习的五个数据集。然后,我们进行了一个新的调查之间的相互作用的数据集特征和性能results.Measurements分类精度和面积下的受试者工作特征(ROC)曲线为每种算法在不同的样本量。主动学习算法的性能进行了比较,被动学习使用成对差异的加权平均值。为了确定为什么不同的数据集上的性能不同,我们测量了每个数据集的多样性和不确定性,使用相对熵和相关的结果与性能difference.Results的DIST和CMB算法的性能优于被动学习。在统计显著性水平设定为0.05的情况下,DIST在所有五个数据集中的表现优于被动学习,而CMB在四个数据集中的表现优于被动学习。我们发现,数据集的多样性和DIV性能之间的强相关性,以及数据集的不确定性和DIST algorithm.Conclusion的性能,对于医疗文本分类,适当的主动学习算法可以产生性能相媲美的被动学习与相当小的训练集。特别是,我们的研究结果表明,DIV表现更好的数据具有较高的多样性和DIST的数据具有较低的不确定性。
Objective This study explores active learning algorithms as a way to reduce the requirements for large training sets in medical text classification tasks.Design Three existing active learning algorithms (distance-based (DIST), diversity-based (DIV), and a combination of both (CMB)) were used to classify text from five datasets. The performance of these algorithms was compared to that of passive learning on the five datasets. We then conducted a novel investigation of the interaction between dataset characteristics and the performance results.Measurements Classification accuracy and area under receiver operating characteristics (ROC) curves for each algorithm at different sample sizes were generated. The performance of active learning algorithms was compared with that of passive learning using a weighted mean of paired differences. To determine why the performance varies on different datasets, we measured the diversity and uncertainty of each dataset using relative entropy and correlated the results with the performance differences.Results The DIST and CMB algorithms performed better than passive learning. With a statistical significance level set at 0.05, DIST outperformed passive learning in all five datasets, while CMB was found to be better than passive learning in four datasets. We found strong correlations between the dataset diversity and the DIV performance, as well as the dataset uncertainty and the performance of the DIST algorithm.Conclusion For medical text classification, appropriate active learning algorithms can yield performance comparable to that of passive learning with considerably smaller training sets. In particular, our results suggest that DIV performs better on data with higher diversity and DIST on data with lower uncertainty.