A simple procedure for directly obtaining haplotype sequences of diploid genomes.

A simple procedure for directly obtaining haplotype sequences of diploid genomes.
复制标题

DOI:
10.1186/s12864-015-1818-4
复制
发表时间:
2015-08-28
期刊:
影响因子:
4.4
通讯作者:
Hall N
Hall N
中科院分区:
生物学2区
文献类型:
--
作者:
Noyes HA;Daly D;Goodhead I;Kay S;Kemp SJ;Kenny J;Saccheri I;Schnabel RD;Taylor JF;Hall N

文献摘要

相似文献

几乎所有基因组测序项目都忽略了二倍体生物体包含两个基因组拷贝的事实,因此发表的内容是两者的复合体。这意味着两个或更多个连锁基因座处的替代等位基因之间的关系丢失。我们开发了一种简化的方法,可以直接从单个生物体中获取每个基因组拷贝的单倍体序列。使用简单的样品制备程序获得三组牛样品的二倍体序列,仅需要显微镜和血细胞计数器。样品为: 1) 来自单头安格斯公牛的淋巴细胞; 2) 来自安格斯公牛的精子细胞; 3) 来自东非瘤牛 (EAZ) 的淋巴细胞,在肯尼亚东部的现场实验室收集和处理。使用从 EAZ 牛的淋巴细胞制备的 fosmid 文库的单倍体序列进行比较。通过在每次稀释时用血细胞计数器计数,将细胞连续稀释至每微升一个细胞的浓度。将一微升样品(每个样品可能含有一个细胞)裂解并分成六等份(未分成等份的精子样品除外)。每个等分试样均用 phi29 聚合酶扩增并测序。通过映射到牛 UMD3.1 参考基因组组装获得重叠群,并通过连接彼此在阈值距离内的相邻重叠群来组装支架。过滤掉似乎含有 CNV 或重复序列的支架,留下 N50 长度为 27-133 kb 且基因组覆盖率为 88-98% 的支架。 SNP 单倍型使用 Single individual Haplotyper 程序进行组装,生成 97-201 kb 的 N50 大小,但基因组覆盖率仅为约 27-68%。该方法可在任何实验室使用,无需特殊设备,成本仅比传统二倍体基因组测序略高。编写了大量用于分析和工作流程管理的软件,并可作为补充数据使用。我们开发了一套实验室方案和软件工具,使任何实验室都能以比传统混合二倍体序列稍高的成本获得单倍型序列。
Almost all genome sequencing projects neglect the fact that diploid organisms contain two genome copies and consequently what is published is a composite of the two. This means that the relationship between alternate alleles at two or more linked loci is lost. We have developed a simplified method of directly obtaining the haploid sequences of each genome copy from an individual organism. The diploid sequences of three groups of cattle samples were obtained using a simple sample preparation procedure requiring only a microscope and a haemocytometer. Samples were: 1) lymphocytes from a single Angus steer; 2) sperm cells from an Angus bull; 3) lymphocytes from East African Zebu (EAZ) cattle collected and processed in a field laboratory in Eastern Kenya. Haploid sequence from a fosmid library prepared from lymphocytes of an EAZ cow was used for comparison. Cells were serially diluted to a concentration of one cell per microlitre by counting with a haemocytometer at each dilution. One microlitre samples, each potentially containing a single cell, were lysed and divided into six aliquots (except for the sperm samples which were not divided into aliquots). Each aliquot was amplified with phi29 polymerase and sequenced. Contigs were obtained by mapping to the bovine UMD3.1 reference genome assembly and scaffolds were assembled by joining adjacent contigs that were within a threshold distance of each other. Scaffolds that appeared to contain artefacts of CNV or repeats were filtered out leaving scaffolds with an N50 length of 27–133 kb and a 88–98 % genome coverage. SNP haplotypes were assembled with the Single Individual Haplotyper program to generate an N50 size of 97–201 kb but only ~27–68 % genome coverage. This method can be used in any laboratory with no special equipment at only slightly higher costs than conventional diploid genome sequencing. A substantial body of software for analysis and workflow management was written and is available as supplementary data. We have developed a set of laboratory protocols and software tools that will enable any laboratory to obtain haplotype sequences at only modestly greater cost than traditional mixed diploid sequences.