Mining New Crystal Protein Genes from Bacillus thuringiensis on the Basis of Mixed Plasmid-Enriched Genome Sequencing and a Computational Pipeline

Mining New Crystal Protein Genes from Bacillus thuringiensis on the Basis of Mixed Plasmid-Enriched Genome Sequencing and a Computational Pipeline
复制标题

DOI:
10.1128/aem.00340-12
复制
发表时间:
2012-04
影响因子:
4.4
通讯作者:
Weixing Ye;Lei Zhu;Yingying Liu;N. Crickmore;Donghai Peng;L. Ruan;Ming Sun
Weixing Ye;Lei Zhu;Yingying Liu;N. Crickmore;Donghai Peng;L. Ruan;Ming Sun
中科院分区:
生物学2区
文献类型:
--
作者:
Weixing Ye;Lei Zhu;Yingying Liu;N. Crickmore;Donghai Peng;L. Ruan;Ming Sun

文献摘要

被引文献

相似文献

摘要我们设计了一个高通量的苏云金芽孢杆菌新晶体蛋白基因(CRY)鉴定系统。该系统的开发有两个目标:(I)利用下一代测序生物技术获得混合质粒富集的苏云金芽胞杆菌基因组序列;(Ii)使用计算流水线(使用BtToxin_scanner)鉴定CrY基因。在我们的流水线方法中,我们使用了三种不同的预测方法:BLAST、隐马尔可夫模型(HMM)和支持向量机(SVM)来预测Cry毒素基因的存在。该流水线被证明是快速的(蛋白质和开放阅读框[ORF]的平均速度为1.02Mb/min,核苷酸序列的平均速度为1.80Mb/min),灵敏(它检测到的蛋白质毒素基因比使用从GenBank下载的基因组序列的关键字提取方法多40%),并且高度特异。从我们实验室收集的21株菌株中,根据它们的质粒型和/或晶体形态进行了选择。从这些菌株中提取富含质粒的基因组DNA,混合后用于Illumina测序。对测序数据进行重新组装,并使用计算流水线鉴定了总共113个候选CRY序列。根据这些候选序列与已知基因序列同源性较低的特点,选择了27个候选序列,并通过聚合酶链式反应获得了8个全长基因。最终确定了苏云金芽胞杆菌毒素命名委员会命名的3个新的冷冻型基因(初级序列)和5个冷冻型基因,它们分别被苏云金芽孢杆菌毒素命名委员会命名为Cry8Ac1、Cry7Ha1、Cry21Ca1、Cry32Fa1和Cry21Da1。该系统既高效又经济,可以极大地加速发现新的水稻抗病基因。
ABSTRACT We have designed a high-throughput system for the identification of novel crystal protein genes (cry) from Bacillus thuringiensis strains. The system was developed with two goals: (i) to acquire the mixed plasmid-enriched genomic sequence of B. thuringiensis using next-generation sequencing biotechnology, and (ii) to identify cry genes with a computational pipeline (using BtToxin_scanner). In our pipeline method, we employed three different kinds of well-developed prediction methods, BLAST, hidden Markov model (HMM), and support vector machine (SVM), to predict the presence of Cry toxin genes. The pipeline proved to be fast (average speed, 1.02 Mb/min for proteins and open reading frames [ORFs] and 1.80 Mb/min for nucleotide sequences), sensitive (it detected 40% more protein toxin genes than a keyword extraction method using genomic sequences downloaded from GenBank), and highly specific. Twenty-one strains from our laboratory's collection were selected based on their plasmid pattern and/or crystal morphology. The plasmid-enriched genomic DNA was extracted from these strains and mixed for Illumina sequencing. The sequencing data were de novo assembled, and a total of 113 candidate cry sequences were identified using the computational pipeline. Twenty-seven candidate sequences were selected on the basis of their low level of sequence identity to known cry genes, and eight full-length genes were obtained with PCR. Finally, three new cry-type genes (primary ranks) and five cry holotypes, which were designated cry8Ac1, cry7Ha1, cry21Ca1, cry32Fa1, and cry21Da1 by the B. thuringiensis Toxin Nomenclature Committee, were identified. The system described here is both efficient and cost-effective and can greatly accelerate the discovery of novel cry genes.