Distinguishing potential bacteria-tumor associations from contamination in a secondary data analysis of public cancer genome sequence data.

Distinguishing potential bacteria-tumor associations from contamination in a secondary data analysis of public cancer genome sequence data.
复制标题

在公共癌基因组序列数据的二级数据分析中,区分潜在的细菌肿瘤关联与污染。

DOI:
10.1186/s40168-016-0224-8
复制
发表时间:
2017-01-25
期刊:
影响因子:
15.5
通讯作者:
Dunning Hotopp JC
Dunning Hotopp JC
中科院分区:
生物学1区
文献类型:
--
作者:
Robinson KM;Crabtree J;Mattick JS;Anderson KE;Dunning Hotopp JC

文献摘要

被引文献

相似文献

已知多种细菌会影响致癌作用。因此,我们试图调查大型公共癌症基因组工作(如癌症基因组图谱 (TCGA))生成的公开可用的全基因组和全转录组测序数据是否可用于识别与癌症相关的细菌。 Burrows-Wheeler 比对器 (BWA) 用于将 TCGA 中的 Illumina 双端测序数据子集与人类参考基因组和 RefSeq 数据库中的所有完整细菌基因组进行比对,以从微生物组中识别细菌读数对。通过仔细考虑所调查的癌症类型中存在的所有细菌类群、它们的相对丰度和批次效应,我们能够从某些类群中识别出一些可能是由污染引起的读数对。特别是,卵巢浆液性囊腺癌(OV)和多形性胶质母细胞瘤(GBM)样本中结核分枝杆菌复合体的存在与样本的测序中心相关。此外,Ralstonia spp 的存在之间存在相关性。以及两个急性髓系白血病 (AML) 样本的特定平板。最后,AML 中的假单胞菌样和不动杆菌样读对之间以及胃腺癌 (STAD) 中的假单胞菌样读对之间仍然存在关联,但无法通过批次效应或系统污染来解释,如其他样本中所见。这种方法表明,可以从公共基因组测序数据中识别人类肿瘤样本中可能存在的细菌,并可以通过进一步的实验进行检查。将来当怀疑细菌与疾病相关时,应更加重视这种方法。本文的在线版本 (doi:10.1186/s40168-016-0224-8) 包含补充材料,可供授权用户使用。
A variety of bacteria are known to influence carcinogenesis. Therefore, we sought to investigate if publicly available whole genome and whole transcriptome sequencing data generated by large public cancer genome efforts, like The Cancer Genome Atlas (TCGA), could be used to identify bacteria associated with cancer. The Burrows-Wheeler aligner (BWA) was used to align a subset of Illumina paired-end sequencing data from TCGA to the human reference genome and all complete bacterial genomes in the RefSeq database in an effort to identify bacterial read pairs from the microbiome. Through careful consideration of all of the bacterial taxa present in the cancer types investigated, their relative abundance, and batch effects, we were able to identify some read pairs from certain taxa as likely resulting from contamination. In particular, the presence of Mycobacterium tuberculosis complex in the ovarian serous cystadenocarcinoma (OV) and glioblastoma multiforme (GBM) samples was correlated with the sequencing center of the samples. Additionally, there was a correlation between the presence of Ralstonia spp. and two specific plates of acute myeloid leukemia (AML) samples. At the end, associations remained between Pseudomonas-like and Acinetobacter-like read pairs in AML, and Pseudomonas-like read pairs in stomach adenocarcinoma (STAD) that could not be explained through batch effects or systematic contamination as seen in other samples. This approach suggests that it is possible to identify bacteria that may be present in human tumor samples from public genome sequencing data that can be examined further experimentally. More weight should be given to this approach in the future when bacterial associations with diseases are suspected. The online version of this article (doi:10.1186/s40168-016-0224-8) contains supplementary material, which is available to authorized users.