CAREER: Genomic Data Science: From Informational Limits to Efficient Algorithms
CAREER: Genomic Data Science: From Informational Limits to Efficient Algorithms
批准号:
2046991
负责人:
Ilan Shomorony
金额:
$50.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-06-01 至 2026-05-31
中文摘要
DNA测序技术的进步为生物和医学科学的革命铺平了道路。通过对人类基因组进行测序,人们可以了解几种疾病的遗传基础,并利用这些信息来开发治疗和预防性护理。通过对病毒和细菌的基因组进行测序,人们可以获得对传染病机制的关键见解。但对大量基因组数据的获取、处理和分析提出了几个基本问题,例如:(I)需要收集多少测序数据才能可靠地了解物种的基因组?(2)在保持其有用性的同时,可以将基因组测序数据压缩到什么程度?(Iii)测序错误如何影响进行生物有效推理的能力?这个项目的目标是开发一个框架,以确定基因组数据科学问题的信息限制,即确定基因组数据可以和不能揭示什么。这将导致以信息最优化的方式处理基因组数据的计算高效算法的发展。该项目还将指导代表人数不足的学生,并为他们提供基因组学领域的研究机会。这些研究工作将直接影响数据科学本科课程的内容,进而产生材料(课堂讲稿、教学视频、数据作业、开放源码),用于向社区传播有关基因组技术的可靠信息。基因组数据科学的一个独特方面是,它主要对序列数据进行操作。拟议的研究将按照三个主要方向进行组织,每个主要方向都集中在处理基因组序列数据时出现的关键数据科学任务:(1)对序列对进行比对,(2)从噪声片段中重建序列,以及(3)基于适当的度量对序列进行聚类。由于大量噪声序列的成对比对通常是基因组数据科学中的瓶颈,因此第一推力将研究如何将这些序列的低维表示或草图最佳地用于比对计算。将利用信源编码框架来研究草图大小和比对计算中引起的失真之间的权衡。第二个推力是关于序列重建,它将解决这样一个事实,即计算复杂性障碍,如NP-硬度,往往不能适当地捕获现实世界问题实例的复杂性。引入了基于实例的信息硬度的概念,以允许开发具有特定于实例的理论保证的高效算法。最后,第三个重点是在元基因组测序的背景下研究序列聚类的问题,目标是确定哪些序列来自相同的微生物基因组。元基因组测序数据聚类的信息论度量将被引入,并用于寻求在数据允许的最大分辨率下解析微生物群落的算法。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Advances in DNA sequencing technologies have paved the way for a revolution in the biological and medical sciences. By sequencing the human genome, one can learn about the genetic basis of several diseases and use this information to develop treatments and preventative care. By sequencing the genomes of viruses and bacteria, one can obtain key insights into the mechanisms of infectious diseases. But the acquisition, processing, and analysis of large amounts of genomic data pose several fundamental questions such as: (i) how much sequencing data need be collected to reliably learn the genome of a species? (ii) how much can one compress genomic sequencing data while maintaining its usefulness? (iii) how do sequencing errors impact the ability to perform biologically valid inference? The goal of this project is to develop a framework to establish the informational limits of genomic data science problems, that is, establish what genomic data can and cannot reveal. This will lead to the development of computationally efficient algorithms that process genomic data in an information-optimal way. The project will also mentor underrepresented students and provide them with research opportunities in the field of genomics. The research efforts will directly shape the contents of an undergraduate course on data science, which, in turn, will produce materials (lecture notes, educational videos, data assignments, open-source code) that will be used to disseminate reliable information about genomic technologies to the community. A distinctive aspect of genomic data science is that it operates mainly on sequence data. The proposed research will be organized along three main thrusts, each one focused on a key data science task that arises when dealing with genomic sequence data: (1) aligning pairs of sequences, (2) reconstructing sequences from noisy fragments, and (3) clustering sequences based on appropriate metrics. Since the pairwise alignment of a large number of noisy sequences is often a bottleneck in genomic data science, the first thrust will study how low-dimensional representations of these sequences, or sketches, can be optimally used for alignment computation. A source-coding framework will be leveraged to study the tradeoffs between sketch size and the incurred distortion in alignment computation. The second thrust, on sequence reconstruction, will tackle the fact that computational complexity obstacles such as NP-hardness often do not appropriately capture the complexity of real-world problem instances. A notion of instance-based informational hardness is introduced to allow the development of efficient algorithms with instance-specific theoretical guarantees. Finally, the third thrust studies the problem of clustering sequences in the context of metagenomic sequencing, where the goal is to determine which sequences come from the same microbial genome. Information-theoretic metrics for the clustering of metagenomic sequencing data will be introduced and used in algorithms that seek to resolve microbial communities at the maximum resolution allowed by the data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1109/isit54713.2023.10206707
发表时间:
2023-05
期刊:
2023 IEEE International Symposium on Information Theory (ISIT)
影响因子:
--
作者:
[Kelly Levick;Ilan Shomorony]
通讯作者:
Kelly Levick;Ilan Shomorony
DOI:
--
发表时间:
2023
期刊:
影响因子:
--
作者:
[Seiyun Shin;Hanfang Zhao;Ilan Shomorony]
通讯作者:
Seiyun Shin;Hanfang Zhao;Ilan Shomorony
Coded Shotgun Sequencing
编码鸟枪测序
DOI:
10.1109/jsait.2022.3151737
发表时间:
2022
期刊:
IEEE Journal on Selected Areas in Information Theory
影响因子:
--
作者:
[Ravi, Aditya Narayan, Vahid, Alireza, Shomorony, Ilan]
通讯作者:
Shomorony, Ilan
Torn-Paper Coding
撕纸编码
DOI:
10.1109/tit.2021.3120920
发表时间:
2021
期刊:
IEEE Transactions on Information Theory
影响因子:
2.5
作者:
[Shomorony, Ilan, Vahid, Alireza]
通讯作者:
Vahid, Alireza
Finding a Burst of Positives via Nonadaptive Semiquantitative Group Testing
通过非适应性半定量群体测试发现积极的爆发
DOI:
--
发表时间:
2023
期刊:
IEEE International Symposium on Information Theory
影响因子:
--
作者:
[Li, Yun-Han, Gabrys, Ryan, Sima, Jin, Shomorony, Ilan, Milenkovic, Olgica]
通讯作者:
Milenkovic, Olgica
共 13 条
CIF: Small: Fundamental Limits of DNA-Based Storage
-
批准号:2007597
-
项目类别:Standard Grant
-
资助金额:$48.41万
-
财政年份:2020
-
负责人:Ilan Shomorony
-
依托单位:
海外基金