课题基金 / 基金详情

Collaborative Research: EAGER: Solving the bait learning problem for large-scale DNA enrichment

Collaborative Research: EAGER: Solving the bait learning problem for large-scale DNA enrichment
合作研究:EAGER:解决大规模 DNA 富集的诱饵学习问题
批准号:
2118252
负责人:
Noelle Noyes
金额:
$2.1万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-09-01 至 2023-08-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
微生物组是指生物样本中的所有微生物,并与许多生物学活动和表型有关。例如,“肠道微生物群”通常指的是在人类消化系统中发现的细菌,它被归因于无数的表型和疾病,包括肥胖、阿尔茨海默病、自闭症谱系障碍和癌症类型。同样,土壤微生物群也被归因于植物的耐旱性、开花时间和抗药性。全面研究生物样本中的微生物是具有挑战性的,因为据估计,其中95%至99%的微生物无法在自然环境之外生活,因此无法在实验室环境中分离。幸运的是,鸟枪式元基因组学可以解决这一挑战,因为它能够将与样本中所有微生物对应的DNA作为输入,并产生与它们对应的DNA串(称为“序列读取”)。然后对这些测序仪读数进行进一步分析,以识别和研究微生物。然而,科学家往往只对有限数量的微生物感兴趣,而不是所有微生物。例如,在研究呼吸道拭子中的微生物的情况下,可能只有新冠肺炎对应的序列才是感兴趣的。猎枪元基因组学将对拭子上发现的所有DNA进行序列读取。幸运的是,有一些实验室方法能够将测序限制在选定的一组DNA序列上,这被称为DNA浓缩。DNA富集会产生一组短的、合成的DNA片段(称为“诱饵”),并将其应用于生物样本,然后这些样本只与DNA的选定部分结合。然后,未结合的DNA被冲洗掉,只留下结合的DNA,然后进行测序。因此,这个过程的第一步是识别诱饵的信息学问题,这些诱饵将丰富一组给定的DNA序列。使信息学问题具有挑战性的两件事是,诱饵的数量应该是最小的,并且诱饵不仅准确地结合到单个DNA序列上,而且它们可以结合到许多允许错配的序列上。这个项目将设计我们将结合在信息学,机器学习和数据集成技术的技术,创造实用的方法来创造诱饵。虽然现有的启发式方法证明了DNA浓缩的实用性,但仍然没有任何方法能够有效地对大型DNA数据库做到这一点。因此,这个项目的重点之一就是创建可伸缩的方法。为此,我们将开发解决两个不同信息学问题的方法:(1)诱饵最小化问题,其目的是确定对整个DNA序列集进行丰富的最小大小的诱饵集;(2)DNA序列最大化问题,其目的是选择可通过使用受限大小的诱饵集来丰富的DNA序列的最大子集。在这里,我们将把生物过程的信息整合到问题公式中,并使用最先进的信息学方法解决问题。该项目将带来几项创新,将对信息学、数据科学和其他科学学科产生重大影响。这项工作的更广泛的影响将包括促进我们的知识,微生物和在弗罗里达大学为从事工程的女孩创建课程。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The microbiome refers to all micro-organisms within a biological sample and has been linked to numerous biological activities and phenotypes. For example, the “gut microbiome”, which commonly refers to the bacteria found within the digestive system in humans, has been attributed to countless phenotypes and diseases, including obesity, Alzheimer’s disease, autism spectrum disorder, and types of cancer. Similarly, the soil microbiome has been attributed to drought tolerance, flowering time, and pesticide resistance in plants. Comprehensively studying the micro-organisms in a biological sample is challenging as that it is estimated that between 95% and 99% of them cannot live outside their natural environments, and therefore, cannot be isolated in a laboratory environment. Fortunately, shotgun metagenomics can address this challenge since it is able to take as input the DNA corresponding to all the micro-organisms within a sample and produce the DNA strings (which are called “sequence reads”) corresponding to them. These sequencer reads are then further analyzed to identify and study the micro-organisms. However, frequently scientists are interested in studying a limited number of micro-organisms rather than all of them. For example, in the case of studying the micro-organisms from a respiratory swab, it may be the case that only the sequence reads corresponding to COVID-19 are of interest. Shotgun metagenomics will produce sequence reads for all DNA found on the swab. Fortunately, there are laboratory methods capable of restricting the sequencing to only a selected set of DNA sequences, which is referred to as DNA enrichment. DNA enrichment creates and applies a set of short, synthetic DNA fragments (called “baits”) to a biological sample which then bind to only selected portions of the DNA. The unbound DNA is then rinsed away, leaving only the bound DNA that is then sequenced. Hence, the first step of this process is the informatics problem of identifying the baits that will enrich for a given set of DNA sequences. Two things that make the informatics problem challenging is that the number of baits should be of minimum size, and the baits do not only bind exactly to a single DNA sequence, but they can bind to many sequences with some allowable mismatches. This project will devise we will combine techniques in informatics, machine learning and data integration techniques to create practical methods for creating baits. While existing heuristics demonstrated the utility of DNA enrichment, there still does not exist any methods that are able to do this efficiently for large DNA databases. Hence, one of the focuses of this project is to create scalable methods. To accomplish this, we will develop methods for solving two different informatics problems: (1) the Bait Minimization problem that aims to identify the set of baits of minimal size that enrich for the entire set of DNA sequences, and (2) the DNA Sequence Maximization problem which aims to select the largest subset of DNA sequences that can be enriched by using a bait set of restricted size. Here, we will integrate the information of the biological process into the problem formulations and solve the problems using state-of-the-art informatic approaches. This project will result in several innovations that will have major impact in informatics, data science and other scientific disciplines. The broader impact of this work will encompass the furtherance of our knowledge micro-organisms and the creation of curriculum for Girls Engaged in Engineering Days at the University of Florida.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)