课题基金 / 基金详情

Leveraging k-mer sketching statistics to enhance metagenomic methods and alignment algorithms

Leveraging k-mer sketching statistics to enhance metagenomic methods and alignment algorithms
利用 k-mer 草图统计来增强宏基因组方法和比对算法
批准号:
10675449
负责人:
Antonio Blanca Pimentel
金额:
$44.35万
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-02 至 2027-05-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
项目摘要 面对日益增长的数据量,MinHash草图绘制及其筛选等草图绘制技术 版本是在促进可伸缩性分析方面最有效的版本之一。尽管如此,生物信息学 使用这些技术的算法没有考虑到素描过程中固有的随机性 以及在产生数据的突变过程中(例如,测序错误或进化突变)。这 Project直接解决了这一限制,为如何绘制这些草图奠定了统计基础 方法与突变过程和基于k-mer的技术相互作用,从而产生了新的算法 重要的生物医学问题。目标1首次推导出以下各项的置信度和预测区间 常用的基于草图的生物信息学数量,到目前为止仅作为点估计存在。 要做到这一点,它依赖于概率论中的复杂技术。目标1奠定的数学基础 不仅将帮助我们实现这一提议的生物学目标,而且还将作为量化的基础 未来基于草图的生物信息学算法的性能。然后,目标2将使用这些结果来 开发第一个元基因组分类描述算法,解决以下情况下存在的不确定性 预测样本中微生物的存在和相对丰度。这将解决一个 通过为研究人员提供一种明智的方法来过滤他们的噪声数据,而不是 牺牲敏感性,从而促进生物医学发现(例如,新的CRISPR系统)。此外,这一点 AIM将产生第一个可扩展的方法,以快速估计非基因组样本的比例 由当前参考数据库描述,从而说明哪些数据集包含最高数量的 新的遗传物质,因此生物发现的可能性(例如,新的抗生素)。目标2将是 使用压缩传感和概率论的技术实现的。AIM 3将同时使用和 扩展目标1的结果以量化地改进计算中最基本的工具之一 生物学家的工具包:序列比对。这将为现代序列比对器配备急需的 显著性分数和置信度区间,以及允许自动选择参数设置 达到预期的精确度或查全率。由于它们在生物医学研究中无处不在,即使是很小的进步 在校准器的精度和功能上都会有极大的影响。目标3将通过以下方式实现 来自概率算法的技术。最后,这项提议的长期目标是提供 研究人员开发了一个工具包,使基于k-mer的可伸缩草图算法的开发无需 牺牲了他们量化统计意义的能力。
英文摘要
Project Summary In the face of increasing data sizes, sketching techniques such as MinHash sketching and its winnowed version have been among the most effective in facilitating scalabile analysis. Frequently though, bioinformatic algorithms using these techniques do not account for the randomness inherent in both the sketching process and in the mutation processes that generate the data (e.g. sequencing errors or evolutionary mutations). This project directly addresses this limitation by laying the statistical foundations for how these sketching approaches interact with mutation processes and k-mer based techniques, resulting in new algorithms for important biomedical problems. Aim 1 derives, for the first time, confidence and prediction intervals for frequently utilized sketching-based bioinformatics quantities that until now existed only as point estimates.To do so, it relies on sophisticated techniques from probability theory. The mathematical foundations laid by Aim 1 will not only help us achieve the biological aims of this proposal, but will also serve as a basis for quantifying the performance of future sketching-based bioinformatics algorithms. Aim 2 will then use these results to develop the first metagenomic taxonomic profiling algorithm that accounts for the uncertainty present when predicting the presence and relative abundance of microorganisms in a sample. This will resolve a long-standing issue in this field by providing researchers an informed way to filter their noisy data without sacrificing sensitivity, thereby facilitating biomedical discoveries (e.g. novel CRISPR systems). In addition, this aim will result in the first scalable method to quickly estimate the fraction of a metagenomic sample that is not described by current reference databases, thus illuminating which datasets contain the highest quantity of novel genetic material and hence possibility for biological discovery (e.g. novel antibiotics). Aim 2 will be achieved using techniques from compressive sensing as well as probability theory. Aim 3 will both use and extend the results of Aim 1 to quantifiably improve one of the most fundamental tools in a computational biologist’s toolkit: sequence alignment. This will equip modern sequence aligners with much needed significance scores and confidence intervals, as well as allow for the automatic selection of parameter settings to achieve a desired precision or recall. Due to their ubiquity in biomedical research, even a small improvement in the accuracy and features of an aligner will have tremendous impact. Aim 3 will be achieved using techniques from probabilistic algorithms. Finally, the long-term objective of this proposal is to provide researchers a toolkit that enables the development of scalable k-mer-based sketching algorithms without sacrificing their ability to quantify statistical significance.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1101/gr.277651.123
发表时间: 2023-07
期刊: GENOME RESEARCH
影响因子: 7
作者: [Rahman Hera, Mahmudur, Pierce-Ward, N Tessa, Koslicki, David]
通讯作者: Koslicki, David
DOI: 10.1093/bioinformatics/btad238
发表时间: 2023-06-30
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者: []
通讯作者:
Finding phylogeny-aware and biologically meaningful averages of metagenomic samples: L 2 UniFrac.
寻找宏基因组样本的系统发育感知和生物学意义的平均值:L 2 UniFrac。
DOI: 10.1101/2023.02.02.526854
发表时间: 2023
期刊: bioRxiv : the preprint server for biology
影响因子: --
作者: [Wei,Wei, Millward,Andrew, Koslicki,David]
通讯作者: Koslicki,David
海外基金