Leveraging k-mer sketching statistics to enhance metagenomic methods and alignment algorithms
Leveraging k-mer sketching statistics to enhance metagenomic methods and alignment algorithms
批准号:
10675449
负责人:
Antonio Blanca Pimentel
金额:
$44.35万
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-02 至 2027-05-31
关键词:
AddressAffectAlgorithmsAntibioticsAreaBioinformaticsBiologicalBiomedical ResearchClustered Regularly Interspaced Short Palindromic RepeatsCommunicationCommunitiesComputational BiologyComputational TechniqueConfidence IntervalsDataData SetDatabasesDevelopmentDimensionsEndowmentEnsureFoundationsFrequenciesFutureGenerationsGenetic MaterialsGoalsMathematicsMeasuresMetagenomicsMethodsModernizationMolecular EvolutionMutateMutationOrganismOutcomeOutputPerformanceProbability TheoryProcessResearch PersonnelSamplingSequence AlignmentSequence Read ArchiveSystemTaxonomyTechniquesTestingTimeUncertaintyVariantWorkcostdriving forceimprovedinnovationinsightmetagenomemicroorganismnovelstatisticstheoriestool
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Project Summary
In the face of increasing data sizes, sketching techniques such as MinHash sketching and its winnowed
version have been among the most effective in facilitating scalabile analysis. Frequently though, bioinformatic
algorithms using these techniques do not account for the randomness inherent in both the sketching process
and in the mutation processes that generate the data (e.g. sequencing errors or evolutionary mutations). This
project directly addresses this limitation by laying the statistical foundations for how these sketching
approaches interact with mutation processes and k-mer based techniques, resulting in new algorithms for
important biomedical problems. Aim 1 derives, for the first time, confidence and prediction intervals for
frequently utilized sketching-based bioinformatics quantities that until now existed only as point estimates.To
do so, it relies on sophisticated techniques from probability theory. The mathematical foundations laid by Aim 1
will not only help us achieve the biological aims of this proposal, but will also serve as a basis for quantifying
the performance of future sketching-based bioinformatics algorithms. Aim 2 will then use these results to
develop the first metagenomic taxonomic profiling algorithm that accounts for the uncertainty present when
predicting the presence and relative abundance of microorganisms in a sample. This will resolve a
long-standing issue in this field by providing researchers an informed way to filter their noisy data without
sacrificing sensitivity, thereby facilitating biomedical discoveries (e.g. novel CRISPR systems). In addition, this
aim will result in the first scalable method to quickly estimate the fraction of a metagenomic sample that is not
described by current reference databases, thus illuminating which datasets contain the highest quantity of
novel genetic material and hence possibility for biological discovery (e.g. novel antibiotics). Aim 2 will be
achieved using techniques from compressive sensing as well as probability theory. Aim 3 will both use and
extend the results of Aim 1 to quantifiably improve one of the most fundamental tools in a computational
biologist’s toolkit: sequence alignment. This will equip modern sequence aligners with much needed
significance scores and confidence intervals, as well as allow for the automatic selection of parameter settings
to achieve a desired precision or recall. Due to their ubiquity in biomedical research, even a small improvement
in the accuracy and features of an aligner will have tremendous impact. Aim 3 will be achieved using
techniques from probabilistic algorithms. Finally, the long-term objective of this proposal is to provide
researchers a toolkit that enables the development of scalable k-mer-based sketching algorithms without
sacrificing their ability to quantify statistical significance.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1101/gr.277651.123
发表时间:
2023-07
期刊:
GENOME RESEARCH
影响因子:
7
作者:
[Rahman Hera, Mahmudur, Pierce-Ward, N Tessa, Koslicki, David]
通讯作者:
Koslicki, David
Finding phylogeny-aware and biologically meaningful averages of metagenomic samples: L 2 UniFrac.
寻找宏基因组样本的系统发育感知和生物学意义的平均值:L 2 UniFrac。
DOI:
10.1101/2023.02.02.526854
发表时间:
2023
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
作者:
[Wei,Wei, Millward,Andrew, Koslicki,David]
通讯作者:
Koslicki,David
DOI:
10.1093/bioinformatics/btad238
发表时间:
2023-06-30
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
[]
通讯作者:
海外基金