Automated Classification of Circulating Tumor Cells and the Impact of Interobsever Variability on Classifier Training and Performance

Automated Classification of Circulating Tumor Cells and the Impact of Interobsever Variability on Classifier Training and Performance
复制标题

DOI:
10.1155/2015/573165
复制
发表时间:
2015-01-01
影响因子:
4.1
通讯作者:
Figge, Marc Thilo
Figge, Marc Thilo
中科院分区:
医学3区
文献类型:
--
作者:
Svensson, Carl-Magnus;Huebler, Ron;Figge, Marc Thilo

文献摘要

被引文献

相似文献

个性化医疗的应用需要整合不同的数据,以确定每个患者独特的临床体质。医疗数据的自动化分析是一个不断发展的领域,其中使用不同的机器学习技术来最大限度地减少手动分析的耗时任务。自动分类器的评估和通常的训练需要手动标记数据作为基础事实。在许多情况下,这种标记并不完美,或者是因为即使对训练有素的专家来说数据也是模糊的,或者是因为错误。在这里,我们研究了观察者之间的变化,包括荧光染色的循环肿瘤细胞的图像数据和它的效果上的性能的两个自动分类器,一个随机森林和支持向量机。我们发现,观察员之间的注释的不确定性限制了自动分类器的性能,特别是当它被包括在测试集上的分类器性能进行测量。随机森林分类器对训练数据中的不确定性具有弹性,而支持向量机的性能高度依赖于训练数据中的不确定性。最后,我们引入了共识数据集作为评估自动分类器的一种可能的解决方案,以最大限度地减少观察者间差异的惩罚。
Application of personalized medicine requires integration of different data to determine each patient's unique clinical constitution. The automated analysis of medical data is a growing field where different machine learning techniques are used to minimize the time-consuming task of manual analysis. The evaluation, and often training, of automated classifiers requires manually labelled data as ground truth. In many cases such labelling is not perfect, either because of the data being ambiguous even for a trained expert or because of mistakes. Here we investigated the interobserver variability of image data comprising fluorescently stained circulating tumor cells and its effect on the performance of two automated classifiers, a random forest and a support vector machine. We found that uncertainty in annotation between observers limited the performance of the automated classifiers, especially when it was included in the test set on which classifier performance was measured. The random forest classifier turned out to be resilient to uncertainty in the training data while the support vector machine's performance is highly dependent on the amount of uncertainty in the training data. We finally introduced the consensus data set as a possible solution for evaluation of automated classifiers that minimizes the penalty of interobserver variability.