PSR: polymorphic SSR retrieval.

PSR: polymorphic SSR retrieval.
复制标题

DOI:
10.1186/s13104-015-1474-4
复制
发表时间:
2015-10-01
期刊:
影响因子:
1.8
通讯作者:
D'Agostino N
D'Agostino N
中科院分区:
其他
文献类型:
--
作者:
Cantarella C;D'Agostino N

文献摘要

被引文献

相似文献

随着高通量测序技术的出现,微卫星的大规模鉴定变得负担得起,特别是针对非模式物种。相比之下,利用序列冗余来自动识别多态微卫星的研究很少。迄今为止,能够管理大量序列数据和处理SAM/BAM文件格式的微卫星重复序列基因分型工具很少。它们大多是为具有高质量参考基因组的人类或模式生物开发并在其上进行测试的。在本文中,我们描述了多态性SSR检索(PSR),一个读取计数器和简单序列重复(SSR)长度多态性检测工具。它是用Perl编写的,用于利用下一代测序(NGS)数据识别完美微卫星的长度多态性。PSR的开发考虑到了植物非模式物种,对于这些物种,从头转录组组装通常是可用于ssr挖掘的第一个序列资源。PSR分为两个模块:读取计数模块(PSR_read_retrieval)识别覆盖完美微卫星全长的所有读取;比较模块(PSR_poly_finder)在每个微卫星位点上检测所有基因型的杂合和纯合等位基因。调用长度多态性和减少误报数量的两个阈值可以由用户定义:重叠重复拉伸的最小读取数和最小读取深度。第一个参数决定是否对含微卫星序列进行处理,第二个参数对次要等位基因的识别起决定性作用。PSR在两个不同的案例研究中进行了测试。第一项研究旨在鉴定两种不同植物基因型的rna测序定义的一组从头组装转录本中的多态性SSRs。第二项研究活动旨在研究新测序的叶绿体基因组集合中的序列变异。在这两种情况下,PSR结果与毛细管凝胶分离得到的结果一致。PSR是专门为从NGS数据中自动化多态微卫星的基因和全基因组鉴定的需要而开发的。它克服了基于前ngs时代开发的工具的现有和耗时的努力的限制。
With the advent of high-throughput sequencing technologies large-scale identification of microsatellites became affordable and was especially directed to non-model species. By contrast, few efforts have been published toward the automatic identification of polymorphic microsatellites by exploiting sequence redundancy. Few tools for genotyping microsatellite repeats have been implemented so far that are able to manage huge amount of sequence data and handle the SAM/BAM file format. Most of them have been developed for and tested on human or model organisms with high quality reference genomes. In this note we describe polymorphic SSR retrieval (PSR), a read counter and simple sequence repeat (SSR) length polymorphism detection tool. It is written in Perl and was developed to identify length polymorphisms in perfect microsatellites exploiting next generation sequencing (NGS) data. PSR has been developed bearing in mind plant non-model species for which de novo transcriptome assembly is generally the first sequence resource available to be used for SSR-mining. PSR is divided into two modules: the read-counting module (PSR_read_retrieval) identifies all the reads that cover the full-length of perfect microsatellites; the comparative module (PSR_poly_finder) detects both heterozygous and homozygous alleles at each microsatellite locus across all genotypes under investigation. Two threshold values to call a length polymorphism and reduce the number of false positives can be defined by the user: the minimum number of reads overlapping the repetitive stretch and the minimum read depth. The first parameter determines if the microsatellite-containing sequence must be processed or not, while the second one is decisive for the identification of minor alleles. PSR was tested on two different case studies. The first study aims at the identification of polymorphic SSRs in a set of de novo assembled transcripts defined by RNA-sequencing of two different plant genotypes. The second research activity aims to investigate sequence variations within a collection of newly sequenced chloroplast genomes. In both the cases PSR results are in agreement with those obtained by capillary gel separation. PSR has been specifically developed from the need to automate the gene-based and genome-wide identification of polymorphic microsatellites from NGS data. It overcomes the limits related to the existing and time-consuming efforts based on tools developed in the pre-NGS era.