Data-driven noise modeling of digital DNA melting analysis enables prediction of sequence discriminating power.

Data-driven noise modeling of digital DNA melting analysis enables prediction of sequence discriminating power.
复制标题

数字 DNA 熔解分析的数据驱动噪声建模能够预测序列辨别力。

DOI:
10.1093/bioinformatics/btaa1053
复制
发表时间:
2021
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Coleman,ToddP
Coleman,ToddP
中科院分区:
--
文献类型:
--
作者:
Langouche,Lennart;Aralar,April;Sinha,Mridu;Lawrence,ShelleyM;Fraley,StephanieI;Coleman,ToddP

文献摘要

相似文献

快速筛选复杂样本以寻找广泛的核酸靶点(如传染病)的需求仍未得到满足。数字高分辨率熔融(dHRM)是一种新兴技术,有潜力通过实现广泛的、快速的核酸序列鉴定来满足这一需求。在这里,我们着手开发一个计算框架,用于估计dHRM技术对定义的序列分析任务的解析能力。通过从实验生成的dHRM数据集中导出噪声模型,并将这些模型应用于硅预测的熔体曲线,我们能够生成合成dHRM数据集,这些数据集忠实地再现了由样本和机器变量引起的现实世界变化。然后,我们使用这些数据集来确定给定应用程序可能出现的最具挑战性的熔体曲线分类任务,并测试基准分类器的性能。结果该工具箱使他们的硅设计和测试的基础广泛的dHRM筛选分析和最佳分类器的选择。对于筛选常见人类细菌病原体的示例应用,我们表明,在存在实验噪声的情况下,具有最相似序列和熔体曲线的人类病原体仍然可以可靠地识别。此外,我们发现集成方法在此任务中优于全系列分类器,并且在某些情况下能够用单核苷酸分辨率解决熔融曲线。可用性和实施数据和代码可在https://github.com/lenlan/dHRM-noise-modeling.Supplementary上获得。补充数据可在bioinformaticsonline上获得。
MotivationThe need to rapidly screen complex samples for a wide range of nucleic acid targets, like infectious diseases, remains unmet. Digital High-Resolution Melt (dHRM) is an emerging technology with potential to meet this need by accomplishing broad-based, rapid nucleic acid sequence identification. Here, we set out to develop a computational framework for estimating the resolving power of dHRM technology for defined sequence profiling tasks. By deriving noise models from experimentally generated dHRM datasets and applying these toin silicopredicted melt curves, we enable the production of synthetic dHRM datasets that faithfully recapitulate real-world variations arising from sample and machine variables. We then use these datasets to identify the most challenging melt curve classification tasks likely to arise for a given application and test the performance of benchmark classifiers.ResultsThis toolbox enables thein silicodesign and testing of broad-based dHRM screening assays and the selection of optimal classifiers. For an example application of screening common human bacterial pathogens, we show that human pathogens having the most similar sequences and melt curves are still reliably identifiable in the presence of experimental noise. Further, we find that ensemble methods outperform whole series classifiers for this task and are in some cases able to resolve melt curves with single-nucleotide resolution.Availability and implementationData and code available on https://github.com/lenlan/dHRM-noise-modeling.Supplementary informationSupplementary data are available atBioinformaticsonline.