Read-mapping using personalized diploid reference genome for RNA sequencing data reduced bias for detecting allele-specific expression.

Read-mapping using personalized diploid reference genome for RNA sequencing data reduced bias for detecting allele-specific expression.
复制标题

DOI:
10.1109/bibmw.2012.6470225
复制
发表时间:
2012-10
期刊:
IEEE International Conference on Bioinformatics and Biomedicine workshops. IEEE International Conference on Bioinformatics and Biomedicine
影响因子:
--
通讯作者:
Qin Z
Qin Z
中科院分区:
其他
文献类型:
--
作者:
Yuan S;Qin Z

文献摘要

相似文献

下一代测序技术已广泛应用于遗传学和基因组学研究的许多领域。分析NGS数据的一个基本问题是将短测序读数映射回参考基因组。大多数现有的软件包依赖于单个统一的参考基因组,并且不自动考虑遗传变异。另一方面,大部分不正确映射的读数影响NGS实验结果的正确解释。例如,Degner等人表明,从RNA测序数据检测等位基因特异性表达偏向于参考等位基因。在这项研究中,我们开发了一种方法,利用DirectX 11支持的图形处理单元(GPU)的并行计算能力,根据该特定个体的所有已知遗传变异产生个性化的二倍体参考基因组。我们表明,使用这样一个个性化的二倍体参考基因组可以提高定位的准确性,并显着减少对参考等位基因的等位基因特异性表达分析的偏见。我们的方法可以应用于任何个人,基因型信息获得基于阵列的基因分型或重测序。除了参考基因组之外,不需要对比对算法进行额外的改变来进行读段映射,因此可以利用任何现有的读段映射工具并实现改进的读段映射结果。软件程序的C++和GPU计算着色器源代码可在http://code.google.com/p/diploid-mapping/downloads/list上获得。
Next generation sequencing (NGS) technologies have been applied extensively in many areas of genetics and genomics research. A fundamental problem when comes to analyzing NGS data is mapping short sequencing reads back to the reference genome. Most of existing software packages rely on a single uniform reference genome and do not automatically take into the consideration of genetic variants. On the other hand, large proportions of incorrectly mapped reads affect the correct interpretation of the NGS experimental results. As an example, Degner et al. showed that detecting allele-specific expression from RNA sequencing data was biased toward the reference allele. In this study, we developed a method that utilize DirectX 11 enabled graphics processing unit (GPU)’s parallel computing power to produces a personalized diploid reference genome based on all known genetic variants of that particular individual. We show that using such a personalized diploid reference genome can improve mapping accuracy and significantly reduce the bias toward reference allele in allele-specific expression analysis. Our method can be applied to any individual that has genotype information obtained either from array-based genotyping or resequencing. Besides the reference genome, no additional changes to alignment algorithm are needed for performing read mapping therefore one can utilize any of the existing read mapping tools and achieve the improved read mapping result. C++ and GPU compute shader source code of the software program is available at: http://code.google.com/p/diploid-mapping/downloads/list.