Automatic Discrimination of Human Hematopoietic Tumor Cell Lines using a Combination of Imaging Flow Cytometry and Convolutional Neural Network.

Automatic Discrimination of Human Hematopoietic Tumor Cell Lines using a Combination of Imaging Flow Cytometry and Convolutional Neural Network.
复制标题

使用成像流式细胞术和卷积神经网络的组合自动区分人类造血肿瘤细胞系。

DOI:
10.1007/s13577-021-00506-2
复制
发表时间:
2021
期刊:
影响因子:
4.3
通讯作者:
Fujioka T
Fujioka T
中科院分区:
生物学3区
文献类型:
--
作者:
Matsuoka Y;Nakatsuka R;Fujioka T

文献摘要

相似文献

准确检测血液系统异常,如造血肿瘤,对于为患者提供适当的治疗非常重要。血液学家一般根据检查结果进行综合诊断,如全血细胞计数、显微镜观察、细胞学、流式细胞术、荧光原位杂交、聚合酶链式反应和G显带[1]。然而,这种诊断方法既耗时又费力,诊断标准也因血液学家的技能而异。相比之下,卷积神经网络(CNN)的最新发展使得与传统的机器学习方法相比能够获得非常高精度的图像分析[2,3]。使用CNN进行基于图像的细胞分类的概念已经被许多研究人员构思,并有望应用于生命科学和医学领域[4]。在血液学领域,形态信息往往在一定程度上帮助血液学家诊断特定类型的肿瘤和细胞状态。因此,使用CNN可以从造血细胞的形态中检测到发生在造血系统中的异常。然而,准备大量的标记数据(带有细胞类型注释的图像)以构成特定CNN的高质量模型需要大量的工作和成本[5]。当CNN应用于血细胞的自动分类时也是如此。例如,使用传统的细胞旋转法和随后的染色获得的大量标记细胞图像的人工制备是困难的。在这方面,成像流式细胞术(IFC)使我们能够轻松地获得大量标记的单个血细胞图像[6]。因此,我们结合成像、流式细胞仪分析和CNN,根据不同的造血肿瘤细胞系的形态特征,很容易地对它们进行分类。首先,我们使用IFC收集了10种不同的人类造血肿瘤细胞系(包括急性髓系白血病、慢性髓系白血病、B细胞急性淋巴细胞白血病和骨髓瘤;列在表S1中)的亮场图像和两个荧光通道数据(图1a和图1a)。S1)。为了获得高质量的训练数据,通过传统的流式细胞仪门控策略,使用两个荧光通道数据(DRAQ5和抗Annexin V-FITC抗体)排除了碎片和死亡细胞(图1a)。此外,使用IFC软件中预先实现的参数(图1a)排除了靠近屏幕边缘的单元格的图像和散焦的图像。每个细胞图像为48×48像素数据(8位,灰度)(图2)。S1)。用于创建测试数据的细胞图像也以相同的方式进行处理。最后,我们准备了包含9.0×105图像(10组,9×104图像/组)的训练数据和包含3.0×104图像(10组,3.0×103图像/组)的测试数据。从10个细胞系的训练图像中提取明显的特征,并建立模型(保存在https://figsh ARE中)。Com/S/0ff58 4cb07 0cd03 164aa,相关预印本存放在https://www.Biorx IV。Org/Conte NT/44682/10.1101 3v2)通过构建CNN(图1b)来构成自动识别。每个细胞系的分类精度随着用于训练的图像数量的增加而增加(图2a,b)。精确度、召回率和
Accurate detection of blood system abnormalities such as hematopoietic tumors is very important in providing appropriate treatment to patients. Hematologists generally make a comprehensive diagnosis based on results of investigations, such as complete blood count, microscopic observation, cytology, flow cytometry, fluorescent in situ hybridization, polymerase chain reaction, and G-banding [1]. However, such diagnostic methods are time-consuming and laborintensive, and diagnostic criteria vary depending on the skills of hematologists. By contrast, the recent development of Convolutional Neural Networks (CNNs) enables in obtaining very high-precision image analysis compared to conventional machine learning methods [2, 3]. The concept of image-based cell classification using CNNs has already been conceived by many researchers and is also expected to be applied in the fields of life sciences and medicine [4]. In the field of hematology, morphological information often helps hematologists to diagnose specific types of tumor and state of cells to some degree. Therefore, it is expected that abnormalities occurring in the hematopoietic system could be detected from hematopoietic cell morphologies using CNNs. However, preparation of large amounts of labeled data (images with annotation of cell types) to constitute a high-quality model of a particular CNN requires a great deal of effort and cost [5]. This is also the case when CNN is applied to automatic classification of blood cells. For example, it is difficult to manually prepare a large number of labeled cell images obtained using conventional cytospin methods and subsequent staining. In this respect, imaging flow cytometry (IFC) enable us to easily obtain large numbers of labeled single blood cell images [6]. Therefore, we combined imaging flow cytometry analysis and CNN to easily classify different hematopoietic tumor cell-derived cell lines using their morphological features. First, we collected bright-field images and two fluorescent channel data from ten different human hematopoietic tumor cell lines (including acute myeloid leukemia, chronic myeloid leukemia, B-cell acute lymphoblastic leukemia, and myeloma; listed in Table S1) using IFC (Fig. 1 a and Fig. S1). To obtain high-quality training data, debris and dead cells were excluded using two fluorescent channel data (DRAQ5 and anti-Annexin V-FITC antibody) by a conventional flow cytometry gating strategy (Fig. 1 a). In addition, images of cells close to the edges of the screen and images that were out of focus were excluded using the pre-implemented parameters in the IFC software (Fig. 1 a). Each cell image was obtained as 48× 48 pixel data (8-bit, grayscale)(Fig. S1). Images of cells used to create test data were also processed in the same manner. Finally, we prepared training data consisting 9.0× 105 images (ten groups comprising 9× 104 images/group) and test data containing 3.0× 104 images (ten groups comprising 3.0× 103 images/group). Distinctive features were extracted from the training images of ten cell lines, and a model (deposited in https://figsh are. com/s/0ff58 4cb07 0cd03 164aa, and related preprint was deposited in https://www. biorx iv. org/conte nt/10.1101/44682 3v2) for automatic discrimination was constituted by constructing a CNN (Fig. 1 b). The accuracy of classification for each cell line increased corresponding to the increasing number of images for training (Fig. 2 a, b). The precision, recall, and