CaPSID: a bioinformatics platform for computational pathogen sequence identification in human genomes and transcriptomes.

CaPSID: a bioinformatics platform for computational pathogen sequence identification in human genomes and transcriptomes.
复制标题

DOI:
10.1186/1471-2105-13-206
复制
发表时间:
2012-08-17
期刊:
影响因子:
3
通讯作者:
Ferretti V
Ferretti V
中科院分区:
生物学4区
文献类型:
--
作者:
Borozan I;Wilson S;Blanchette P;Laflamme P;Watt SN;Krzyzanowski PM;Sircoulomb F;Rottapel R;Branton PE;Ferretti V

文献摘要

参考文献

被引文献

相似文献

现在已经确定,近20%的人类癌症是由传染性病原体引起的,未来人类致癌病原体的清单将会随着各种癌症类型的增加而增加。利用下一代测序技术进行全肿瘤转录组和基因组测序为人类组织中的病原体检测和发现提供了无与伦比的机会,但这需要开发新的全基因组生物信息学工具。在这里,我们提出了CaPSID(计算病原体序列鉴定),这是一个综合的生物信息学平台,用于鉴定、查询和可视化肿瘤基因组和转录组中外源性和内源性病原体核苷酸序列。CaPSID包括一个可扩展的高性能数据存储数据库和一个集成了基因组浏览器JBrowse的web应用程序。CaPSID还为预对齐BAM文件的序列分析提供了有用的指标,例如基因和基因组覆盖率,并且经过优化,可以在内存使用率低的多处理器计算机上高效运行。为了证明CaPSID的实用性和有效性,我们对模拟数据集和卵巢癌转录组样本进行了全面分析。CaPSID正确识别了模拟数据集中的所有人类和病原体序列,而在卵巢数据集中,CaPSID的预测在体外成功验证。
It is now well established that nearly 20% of human cancers are caused by infectious agents, and the list of human oncogenic pathogens will grow in the future for a variety of cancer types. Whole tumor transcriptome and genome sequencing by next-generation sequencing technologies presents an unparalleled opportunity for pathogen detection and discovery in human tissues but requires development of new genome-wide bioinformatics tools. Here we present CaPSID (Computational Pathogen Sequence IDentification), a comprehensive bioinformatics platform for identifying, querying and visualizing both exogenous and endogenous pathogen nucleotide sequences in tumor genomes and transcriptomes. CaPSID includes a scalable, high performance database for data storage and a web application that integrates the genome browser JBrowse. CaPSID also provides useful metrics for sequence analysis of pre-aligned BAM files, such as gene and genome coverage, and is optimized to run efficiently on multiprocessor computers with low memory usage. To demonstrate the usefulness and efficiency of CaPSID, we carried out a comprehensive analysis of both a simulated dataset and transcriptome samples from ovarian cancer. CaPSID correctly identified all of the human and pathogen sequences in the simulated dataset, while in the ovarian dataset CaPSID’s predictions were successfully validated in vitro.
DOI: 10.1093/nar/gkr948
发表时间: 2012-01
影响因子: 14.9
作者:
Hunter S;Jones P;Mitchell A;Apweiler R;Attwood TK;Bateman A;Bernard T;Binns D;Bork P;Burge S;de Castro E;Coggill P;Corbett M;Das U;Daugherty L;Duquenne L;Finn RD;Fraser M;Gough J;Haft D;Hulo N;Kahn D;Kelly E;Letunic I;Lonsdale D;Lopez R;Madera M;Maslen J;McAnulla C;McDowall J;McMenamin C;Mi H;Mutowo-Muellenet P;Mulder N;Natale D;Orengo C;Pesseat S;Punta M;Quinn AF;Rivoire C;Sangrador-Vegas A;Selengut JD;Sigrist CJ;Scheremetjew M;Tate J;Thimmajanarthanan M;Thomas PD;Wu CH;Yeats C;Yong SY
通讯作者: Yong SY
DOI: 10.1038/nature08987
发表时间: 2010-04-15
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
Biopython:用于计算分子生物学和生物信息学的免费 Python 工具。
DOI: 10.1093/bioinformatics/btp163
发表时间: 2009-06-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Cock PJ;Antao T;Chang JT;Chapman BA;Cox CJ;Dalke A;Friedberg I;Hamelryck T;Kauff F;Wilczynski B;de Hoon MJ
通讯作者: de Hoon MJ
DOI: 10.1128/mcb.24.21.9619-9629.2004
发表时间: 2004-11-01
影响因子: 5.3
作者:
Blanchette, P;Cheng, CY;Branton, PE
通讯作者: Branton, PE
DOI: 10.1101/gr.094607.109
发表时间: 2009-09-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Skinner, Mitchell E.;Uzilov, Andrew V.;Holmes, Ian H.
通讯作者: Holmes, Ian H.