课题基金 / 基金详情

RII Track-4: NSF: Extracting Pan Genomic Information from Metagenomic Data: Distributed Algorithms and Scalable Software

RII Track-4: NSF: Extracting Pan Genomic Information from Metagenomic Data: Distributed Algorithms and Scalable Software
RII Track-4:NSF:从宏基因组数据中提取泛基因组信息:分布式算法和可扩展软件
批准号:
2327456
负责人:
Arghya Das
金额:
$29.21万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-01-01 至 2025-12-31

项目摘要

项目成果

Arghya Das的其他基金

相似基金

相关文献

中文摘要
翻译
对元基因组的分析,即直接从环境样本收集的遗传数据,已成为各种研究领域的组成部分,包括气候研究、人类健康、稀土元素的发现、社会和环境复原力规划等。以这种方式收集的遗传数据通常包含多个微生物群落的混合种群。科学家经常致力于从这样一个混合种群中提取和表示特定微生物物种的基因组多样性信息,这一过程被称为泛基因组信息表示。然而,缺乏理论上合理的和生物学上有效的能够执行元基因组或泛基因组分析的算法。此外,高通量基因组测序机产生的海量遗传数据要求这些算法在自然界中是可扩展和分布式的。该项目将研究从大规模元基因组数据集中提取泛基因组信息的分布式算法方面及其实际实现。这项研究与阿拉斯加EPSCoR在其最新科学和技术计划中优先考虑的至少六个不同研究领域一致,包括社区复原力、资源开采、食物-能源-水联系、可再生资源、环境监测和单一健康。这项RII Track-4:NSF奖学金将使阿拉斯加大学费尔班克斯大学(UAF)的一名助理教授和一名研究生能够与北卡罗来纳州立大学(NCSU)的科学家合作并利用他们的资源。首席调查员(PI)将与生物信息学和算法领域的专家合作,开发一套可证明正确的、可扩展的、低时间复杂度的分布式算法,用于从大规模元基因组数据集中提取泛基因组信息。此外,利用NCSU的尖端高性能计算(HPC)资源,PI旨在创建实现这些算法的HPC兼容软件框架的初步版本。分析流程包括四个不同的阶段:1)元基因组纠错,2)元基因组组装,3)组装基因组的装订和注释,以及4)创建可用微生物的泛基因组图谱。这些阶段中的每一个都带来了算法方面的挑战。元基因组数据集中微生物群的多样性,加上仪器误差,使得识别实际物种及其遗传多样性的过程具有极大的挑战性,需要在字符串匹配和图形分析方面进行广泛的研究。分布式软件实施必须应对众多高性能计算挑战。研究成果,包括出版物和开源代码库,将支持UAF的多项研究活动,重点是北极气候变化、北极海洋生物学、阿拉斯加土著健康等。由该奖学金促成的合作还将为UAF的跨学科博士项目奠定基础,该项目涵盖计算机科学、生物信息学和本土科学专业。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The analysis of metagenomes, i.e., genetic data collected directly from environmental samples, has become integral to various research areas, including climate studies, human health, the discovery of rare earth elements, social and environmental resilience planning, and more. Genetic data collected in this manner typically contains a mixed population of multiple microbial communities. Scientists often aim to extract and represent the genomic diversity information of a particular microbial species from such a mixed population, a process known as pan-genomic information representation. However, there is a shortage of theoretically sound and biologically valid algorithms capable of performing metagenomic or pan-genomic analysis. Furthermore, the vast amount of genetic data generated by high-throughput genome sequencing machines necessitates that these algorithms be scalable and distributed in nature. This project will investigate both the distributed algorithmic aspect and its practical implementation to extract pan-genomic information from large-scale metagenomic datasets. This research aligns with at least six different research areas prioritized by Alaska EPSCoR in their latest Science and Technology Plan, including Community Resilience, Resource Extraction, Food-Energy-Water Nexus, Renewable Resources, Environmental Monitoring, and One Health.This RII Track-4: NSF fellowship will enable an Assistant Professor and a graduate student at the University of Alaska Fairbanks (UAF) to collaborate with scientists at North Carolina State University (NCSU) and utilize their resources. The Principal Investigator (PI) will work alongside experts in the field of bioinformatics and algorithms to develop a set of provably correct, scalable, and distributed algorithms with low time complexity for extracting pan-genomic information from large-scale metagenomic datasets. Additionally, utilizing cutting-edge high-performance computing (HPC) resources at NCSU, the PI aims to create a preliminary version of an HPC-compliant software framework implementing these algorithms. The analytic pipeline comprises four distinct stages: 1) metagenomic error correction, 2) metagenomic assembly, 3) binning and annotation of the assembled genome, and 4) creating the pan-genomic profile of the available microbes. Each of these stages presents algorithmic challenges. The diverse coverages of microbiomes in the metagenomic dataset, coupled with instrumental errors, render the process of identifying the actual species and their genetic diversity exceedingly challenging, necessitating extensive research in string matching and graph analysis. The distributed software implementation must address numerous HPC challenges. The research outcomes, including publications and open-source codebases, will support multiple research activities at UAF, focusing on arctic climate change, arctic marine biology, Alaska Native health, among others. The collaboration facilitated by this fellowship will also lay the foundation for an interdisciplinary Ph.D. program at UAF, encompassing computer science, bioinformatics, and indigenous science concentrations.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Equipment: MRI Track-I: Acquisition of CyBR: Cyber Infrastructure for Big Data Research Critical for Alaska
海外基金