Novel Sequence Discovery by Subtractive Genomics

Novel Sequence Discovery by Subtractive Genomics
复制标题

DOI:
10.3791/58877
复制
发表时间:
2019-01-01
影响因子:
1.2
通讯作者:
Bracht, John R.
Bracht, John R.
中科院分区:
综合性期刊4区
文献类型:
--
作者:
Asalone, Kathryn C.;Nelson, Megan M.;Bracht, John R.

文献摘要

被引文献

相似文献

减法基因组学可以用于任何研究,其目的是确定基因,蛋白质或一般区域的序列,嵌入在更大的基因组环境中。减法基因组学使研究人员能够通过全面测序和减去已知的遗传元素来分离感兴趣的目标序列(T)(参考文献,R)。该方法可用于鉴定线粒体、叶绿体、病毒或种系限制性染色体等新序列,当T难以从R中分离出来时特别有用。该方法从全面的基因组数据(R + T)开始,使用基本局部比对搜索工具(BLAST)对参考序列或序列去除匹配的已知序列(R),留下目标序列(T)。为了使减法效果最好,R应该是一个相对完整的草稿,缺少t。由于减法后剩下的序列是通过定量pcr (quantitative Polymerase Chain Reaction, qPCR)来测试的,因此R不需要是完整的,该方法就可以工作。在这里,我们将计算步骤与实验步骤连接到一个循环中,该循环可以根据需要迭代,依次删除多个参考序列并改进对t的搜索。减去基因组学的优势在于,即使在物理纯化困难、不可能或昂贵的情况下,也可以识别出全新的目标序列。该方法的缺点是寻找合适的减法参考,并获得用于qPCR检测的t阳性和阴性样本。我们描述了从斑胸草雀生殖系限制性染色体中鉴定第一个基因的方法的实现。在这种情况下,计算过滤涉及三个参考(R),在三个周期内依次删除:不完整的基因组组装、原始基因组数据和转录组数据。
Subtractive genomics can be used in any research where the goal is to identify the sequence of a gene, protein, or general region that is embedded in a larger genomic context. Subtractive genomics enables a researcher to isolate a target sequence of interest (T) by comprehensive sequencing and subtracting out known genetic elements (reference, R). The method can be used to identify novel sequences such as mitochondria, chloroplasts, viruses, or germline restricted chromosomes, and is particularly useful when T cannot be easily isolated from R. Beginning with the comprehensive genomic data (R + T), the method uses Basic Local Alignment Search Tool (BLAST) against a reference sequence, or sequences, to remove the matching known sequences (R), leaving behind the target (T). For subtraction to work best, R should be a relatively complete draft that is missing T. Since sequences remaining after subtraction are tested through quantitative Polymerase Chain Reaction (qPCR), R does not need to be complete for the method to work. Here we link computational steps with experimental steps into a cycle that can be iterated as needed, sequentially removing multiple reference sequences and refining the search for T. The advantage of subtractive genomics is that a completely novel target sequence can be identified even in cases in which physical purification is difficult, impossible, or expensive. A drawback of the method is finding a suitable reference for subtraction and obtaining T-positive and negative samples for qPCR testing. We describe our implementation of the method in the identification of the first gene from the germline-restricted chromosome of zebra finch. In that case computational filtering involved three references (R), sequentially removed over three cycles: an incomplete genomic assembly, raw genomic data, and transcriptomic data.