Overlapping pools for high-throughput targeted resequencing

Overlapping pools for high-throughput targeted resequencing
复制标题

DOI:
10.1101/gr.088559.108
复制
发表时间:
2009-07-01
期刊:
影响因子:
7
通讯作者:
Pe'er, Itsik
Pe'er, Itsik
中科院分区:
生物学1区
文献类型:
--
作者:
Prabhu, Snehit;Pe'er, Itsik

文献摘要

被引文献

相似文献

对个体池中的基因组 DNA 进行重新测序是检测目标区域新变异并在病例和对照之间进行比较的有效策略。有多种方法可以将个体分配到要对其进行测序的池中。目前主要使用的简单、不相交的池方案(许多个体到一个池)提供了对等位基因频率的洞察,但不提供等位基因携带者的身份。我们提出了一个重叠池设计的框架,其中每个单独的样本在多个池中重新测序(许多个体到许多池)。一旦发现变体,观察到该变体的一组池就会揭示其携带者的身份。我们形式化了此类池设计的数学框架,并列出了此类设计的要求。我们特别解决了合并重测序设计的三个实际问题:(1)由于扩增和测序过程中引入的错误而导致的假阳性; (2) 由于覆盖不均匀而加剧的特定等位基因采样不足而导致的假阴性;因此,(3) 在存在错误的情况下对各个承运人的识别不明确。我们以纠错码理论为基础来设计克服这些缺陷的池。我们表明,在重测序研究的实际参数中,我们的设计保证了明确的单例载体识别的高概率,同时在敏感性、特异性和估计等位基因频率的能力方面保持了初始池的特征。我们展示了我们的设计使用来自 1000 Genomes Pilot 3 项目的短读数据提取罕见变异的能力。
Resequencing genomic DNA from pools of individuals is an effective strategy to detect new variants in targeted regions and compare them between cases and controls. There are numerous ways to assign individuals to the pools on which they are to be sequenced. The naive, disjoint pooling scheme (many individuals to one pool) in predominant use today offers insight into allele frequencies, but does not offer the identity of an allele carrier. We present a framework for overlapping pool design, where each individual sample is resequenced in several pools (many individuals to many pools). Upon discovering a variant, the set of pools where this variant is observed reveals the identity of its carrier. We formalize the mathematical framework for such pool designs and list the requirements from such designs. We specifically address three practical concerns for pooled resequencing designs: (1) false-positives due to errors introduced during amplification and sequencing; (2) false-negatives due to undersampling particular alleles aggravated by nonuniform coverage; and consequently, (3) ambiguous identification of individual carriers in the presence of errors. We build on theory of error-correcting codes to design pools that overcome these pitfalls. We show that in practical parameters of resequencing studies, our designs guarantee high probability of unambiguous singleton carrier identification while maintaining the features of naive pools in terms of sensitivity, specificity, and the ability to estimate allele frequencies. We demonstrate the ability of our designs in extracting rare variations using short read data from the 1000 Genomes Pilot 3 project.