GBScleanR: Robust genotyping error correction using hidden Markov model with error pattern recognition

GBScleanR: Robust genotyping error correction using hidden Markov model with error pattern recognition
复制标题

GBScleanR:使用具有错误模式识别功能的隐马尔可夫模型进行稳健的基因分型错误校正

DOI:
10.1101/2022.03.18.484886
复制
发表时间:
2022
期刊:
bioRxiv
影响因子:
--
通讯作者:
Ashikari Motoyuki
Ashikari Motoyuki
中科院分区:
--
文献类型:
--
作者:
Furuta Tomoyuki;Yamamoto Toshio;Ashikari Motoyuki

文献摘要

相似文献

简化代表性测序(RRS)提供了具有成本效益和节省时间的基因分型平台。尽管RRS在通量方面具有突出的优势,但获得的基因型数据通常包含大量的错误。已经开发了几种采用隐马尔可夫模型(HMM)的纠错方法来克服这些问题。这些方法假设标记物具有均匀的错误率,在等位基因读段比中没有偏差。然而,由于基因组片段的不均匀扩增和读段错误映射,确实会出现偏差。在本文中,我们介绍了一个错误校正工具,GBScleanR,它使强大的和精确的错误校正噪声RRS为基础的基因型数据,将标记特定的错误率到HMM。结果表明,与模拟数据集中的现有工具相比,GBScleanR最多将准确度提高了25个百分点以上,并且即使使用易错标记,也能在真实的数据中实现最可靠的基因型估计。
Reduced-representation sequencing (RRS) provides cost-effective and time-saving genotyping platforms. Despite the outstanding advantage of RRS in throughput, the obtained genotype data usually contain a large number of errors. Several error correction methods employing the hidden Markov model (HMM) have been developed to overcome these issues. These methods assume that markers have a uniform error rate with no bias in the allele read ratio. However, bias does occur because of uneven amplification of genomic fragments and read mismapping. In this paper, we introduce an error correction tool, GBScleanR, which enables robust and precise error correction for noisy RRS-based genotype data by incorporating marker-specific error rates into the HMM. The results indicate that GBScleanR improves the accuracy by more than 25 percentage points at maximum compared to the existing tools in simulation data sets and achieves the most reliable genotype estimation in real data even with error-prone markers.