Fast automated cell phenotype image classification

Fast automated cell phenotype image classification
复制标题

DOI:
10.1186/1471-2105-8-110
复制
发表时间:
2007-03-30
期刊:
影响因子:
3
通讯作者:
Teasdale, Rohan D.
Teasdale, Rohan D.
中科院分区:
生物学4区
文献类型:
--
作者:
Hamilton, Nicholas A.;Pantelic, Radosav S.;Teasdale, Rohan D.

文献摘要

被引文献

相似文献

背景:基因组革命导致了基因和蛋白质测序的快速增长,现在人们的注意力转向了编码蛋白质的功能。在这方面,蛋白质亚细胞定位的显微镜成像被证明是无价的,自动荧光显微镜的最新进展使蛋白质定位能够以高通量成像。因此,需要大规模的自动计算技术来有效地量化、区分和分类亚细胞图像。虽然图像统计已被证明在区分局部化方面非常成功,但常用的测量方法计算相对较慢,通常需要从实验图像中单独选择细胞,从而限制了吞吐量和潜在应用的范围。这里我们介绍了阈值邻接统计量,其实质是对图像进行阈值处理,并统计与给定数目的阈值以上像素相邻的阈值像素数。结果:将阈值邻接统计量应用于蛋白质亚细胞定位图像的分类,具有快速、准确的特点。它们在两个图像组上进行测试(可下载),其中一个图像组的荧光标记蛋白质在10个亚细胞位置内源性表达,另一个图像组的蛋白质被转染到11个位置。对于每个图像集,训练和测试一个支持向量机。内源性和转染组的分类正确率分别为94.4%和86.6%。研究发现,与其他常用的统计数据相比,阈值邻接统计数据提供了相当或更高的精确度,同时计算速度快了一个数量级。此外,阈值邻接统计量与Haralick度量相结合,对内源集和转染集的准确率分别为98.2%和93.2%。结论:阈值邻接统计量有可能极大地扩展图像统计学在计算图像分析中的应用范围和规模。它们消除了从图像中裁剪单个细胞的需要,并且计算速度比其他常用统计数据快了一个数量级,同时提供了类似或更好的分类精度,这两个基本要求都适用于大规模方法。
Background: The genomic revolution has led to rapid growth in sequencing of genes and proteins, and attention is now turning to the function of the encoded proteins. In this respect, microscope imaging of a protein's sub-cellular localisation is proving invaluable, and recent advances in automated fluorescent microscopy allow protein localisations to be imaged in high throughput. Hence there is a need for large scale automated computational techniques to efficiently quantify, distinguish and classify sub-cellular images. While image statistics have proved highly successful in distinguishing localisation, commonly used measures suffer from being relatively slow to compute, and often require cells to be individually selected from experimental images, thus limiting both throughput and the range of potential applications. Here we introduce threshold adjacency statistics, the essence which is to threshold the image and to count the number of above threshold pixels with a given number of above threshold pixels adjacent. These novel measures are shown to distinguish and classify images of distinct sub-cellular localization with high speed and accuracy without image cropping.Results: Threshold adjacency statistics are applied to classification of protein sub-cellular localization images. They are tested on two image sets (available for download), one for which fluorescently tagged proteins are endogenously expressed in 10 sub-cellular locations, and another for which proteins are transfected into 11 locations. For each image set, a support vector machine was trained and tested. Classification accuracies of 94.4% and 86.6% are obtained on the endogenous and transfected sets, respectively. Threshold adjacency statistics are found to provide comparable or higher accuracy than other commonly used statistics while being an order of magnitude faster to calculate. Further, threshold adjacency statistics in combination with Haralick measures give accuracies of 98.2% and 93.2% on the endogenous and transfected sets, respectively.Conclusion: Threshold adjacency statistics have the potential to greatly extend the scale and range of applications of image statistics in computational image analysis. They remove the need for cropping of individual cells from images, and are an order of magnitude faster to calculate than other commonly used statistics while providing comparable or better classification accuracy, both essential requirements for application to large-scale approaches.