Evaluating the performance of targeted sequence capture, RNA‐Seq, and degenerate‐primer PCR cloning for sequencing the largest mammalian multigene family

Evaluating the performance of targeted sequence capture, RNA‐Seq, and degenerate‐primer PCR cloning for sequencing the largest mammalian multigene family
复制标题

DOI:
10.1111/1755-0998.13093
复制
发表时间:
2020-01
影响因子:
7.7
通讯作者:
Laurel R. Yohe;Kalina T. J. Davies;N. Simmons;K. Sears;E. Dumont;S. Rossiter;Liliana M. Dávalos
Laurel R. Yohe;Kalina T. J. Davies;N. Simmons;K. Sears;E. Dumont;S. Rossiter;Liliana M. Dávalos
中科院分区:
生物学1区
文献类型:
--
作者:
Laurel R. Yohe;Kalina T. J. Davies;N. Simmons;K. Sears;E. Dumont;S. Rossiter;Liliana M. Dávalos

文献摘要

被引文献

相似文献

多基因家族通过重复从单拷贝祖先基因进化而来,通常编码对关键生物过程至关重要的蛋白质。这些基因家族的分子分析需要高置信度的序列,但成员的高序列相似性会给测序和下游分析带来挑战。以普通吸血蝙蝠为研究对象,我们评估了不同的测序方法在恢复最大的哺乳动物蛋白编码多基因家族:嗅觉受体(OR)中的表现。以基因组为参照,我们通过以下方法确定了完整蛋白编码受体的比例:(a)通过Sanger技术测序的简并引物扩增子,(b)主要嗅上皮的RNA - Seq,以及(c)从密切相关物种的转录组设计的探针捕获的基因。我们最初对高质量吸血蝙蝠基因组的重新注释结果是,完整的OR基因达到了40400个,是最初估计的两倍多。Sanger测序扩增子在三种方法中表现最差,检测到50%的注释基因组ORs,而靶向序列捕获恢复了近75%的注释基因。每种测序方法都能组装高质量的序列,即使它不能恢复基因组中的所有受体。虽然一些差异可能是由于研究设计的局限性(例如,不同的个体),但方法之间的差异主要是由于某些受体的低覆盖率而不是高装配错误率引起的。鉴于这种可变性,我们警告不要使用每个物种完整受体的计数来模拟多基因家族的出生-死亡过程。相反,我们的结果支持使用同源序列来探索和模拟塑造这些基因的进化过程。
Multigene families evolve from single‐copy ancestral genes via duplication, and typically encode proteins critical to key biological processes. Molecular analyses of these gene families require high‐confidence sequences, but the high sequence similarity of the members can create challenges for sequencing and downstream analyses. Focusing on the common vampire bat, Desmodus rotundus, we evaluated how different sequencing approaches performed in recovering the largest mammalian protein‐coding multigene family: olfactory receptors (OR). Using the genome as a reference, we determined the proportion of intact protein‐coding receptors recovered by: (a) amplicons from degenerate primers sequenced via Sanger technology, (b) RNA‐Seq of the main olfactory epithelium, and (c) those genes captured with probes designed from transcriptomes of closely‐related species. Our initial re‐annotation of the high‐quality vampire bat genome resulted in >400 intact OR genes, more than doubling the original estimate. Sanger‐sequenced amplicons performed the poorest among the three approaches, detecting 50% of the annotated genomic ORs, and targeted sequence capture recovered nearly 75% of annotated genes. Each sequencing approach assembled high‐quality sequences, even if it did not recover all receptors in the genome. While some variation may be due to limitations of the study design (e.g., different individuals), variation among approaches was mostly caused by low coverage of some receptors rather than high rates of assembly error. Given this variability, we caution against using the counts of intact receptors per species to model the birth‐death process of multigene families. Instead, our results support the use of orthologous sequences to explore and model the evolutionary processes shaping these genes.