课题基金 / 基金详情

Leveraging k-mer sketching statistics to enhance metagenomic methods and alignment algorithms

Leveraging k-mer sketching statistics to enhance metagenomic methods and alignment algorithms
利用 k-mer 草图统计来增强宏基因组方法和比对算法
批准号:
10675449
负责人:
Antonio Blanca Pimentel
金额:
$44.35万
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-02 至 2027-05-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Project Summary In the face of increasing data sizes, sketching techniques such as MinHash sketching and its winnowed version have been among the most effective in facilitating scalabile analysis. Frequently though, bioinformatic algorithms using these techniques do not account for the randomness inherent in both the sketching process and in the mutation processes that generate the data (e.g. sequencing errors or evolutionary mutations). This project directly addresses this limitation by laying the statistical foundations for how these sketching approaches interact with mutation processes and k-mer based techniques, resulting in new algorithms for important biomedical problems. Aim 1 derives, for the first time, confidence and prediction intervals for frequently utilized sketching-based bioinformatics quantities that until now existed only as point estimates.To do so, it relies on sophisticated techniques from probability theory. The mathematical foundations laid by Aim 1 will not only help us achieve the biological aims of this proposal, but will also serve as a basis for quantifying the performance of future sketching-based bioinformatics algorithms. Aim 2 will then use these results to develop the first metagenomic taxonomic profiling algorithm that accounts for the uncertainty present when predicting the presence and relative abundance of microorganisms in a sample. This will resolve a long-standing issue in this field by providing researchers an informed way to filter their noisy data without sacrificing sensitivity, thereby facilitating biomedical discoveries (e.g. novel CRISPR systems). In addition, this aim will result in the first scalable method to quickly estimate the fraction of a metagenomic sample that is not described by current reference databases, thus illuminating which datasets contain the highest quantity of novel genetic material and hence possibility for biological discovery (e.g. novel antibiotics). Aim 2 will be achieved using techniques from compressive sensing as well as probability theory. Aim 3 will both use and extend the results of Aim 1 to quantifiably improve one of the most fundamental tools in a computational biologist’s toolkit: sequence alignment. This will equip modern sequence aligners with much needed significance scores and confidence intervals, as well as allow for the automatic selection of parameter settings to achieve a desired precision or recall. Due to their ubiquity in biomedical research, even a small improvement in the accuracy and features of an aligner will have tremendous impact. Aim 3 will be achieved using techniques from probabilistic algorithms. Finally, the long-term objective of this proposal is to provide researchers a toolkit that enables the development of scalable k-mer-based sketching algorithms without sacrificing their ability to quantify statistical significance.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1101/gr.277651.123
发表时间: 2023-07
期刊: GENOME RESEARCH
影响因子: 7
作者: [Rahman Hera, Mahmudur, Pierce-Ward, N Tessa, Koslicki, David]
通讯作者: Koslicki, David
Finding phylogeny-aware and biologically meaningful averages of metagenomic samples: L 2 UniFrac.
寻找宏基因组样本的系统发育感知和生物学意义的平均值:L 2 UniFrac。
DOI: 10.1101/2023.02.02.526854
发表时间: 2023
期刊: bioRxiv : the preprint server for biology
影响因子: --
作者: [Wei,Wei, Millward,Andrew, Koslicki,David]
通讯作者: Koslicki,David
DOI: 10.1093/bioinformatics/btad238
发表时间: 2023-06-30
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者: []
通讯作者:
海外基金