课题基金 / 基金详情

Collaborative Research: CIF: Small: Coded String Reconstruction Problems in Molecular Storage

Collaborative Research: CIF: Small: Coded String Reconstruction Problems in Molecular Storage
合作研究:CIF:小型:分子存储中的编码串重建问题
批准号:
2008125
负责人:
Olgica Milenkovic
金额:
$23.15万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2024-09-30

项目摘要

项目成果

Olgica Milenkovic的其他基金

相似基金

相关文献

中文摘要
翻译
由于目前的DNA和蛋白质测序平台无法读取表示蛋白质/基因序列的长字符串的内容,因此从它们的片段或大量片段重建文本串的问题在计算生物学中具有重要的意义。串重建的一个典型例子出现在DNA组装中:在DNA组装中,人们创建同一长串的多个副本,并剪切这些副本以读出重叠的短子串,这些子串可以通过匹配它们的前缀和后缀来组合在一起。由于分段和匹配过程中的错误,重建的串可能不是原始串的完美复制品。此外,对于许多字符串来说,唯一重建本质上是不可能的。这是用于基础生物学研究的下一代测序技术的一个主要问题,因为在这种情况下不可能确保明确的结果。该项目中追求的成功代码设计可以解决阻碍新兴分子计算和存储范例实现的可靠性和内容检索问题。这个项目致力于开发新的编码方法,用于根据组成的子串、子序列和子串组成来唯一地重建字符串或字符串池。所采用的技术代表了新的图论、组合优化和信息论方法的组合。特别是,该项目将调查平衡的部分De Bruijn串用于基于子串的重建、类似加泰罗尼亚的路径用于多组合成重建以及编码的多道重建方法,涉及对删除校正码和叠加码的专门修改。编码方案将在伊利诺伊斯大学开发的基于DNA和合成聚合物的数据存储平台上进行测试。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The problem of reconstructing text strings from their fragments or masses of fragments is of focal importance in computational biology as current DNA and protein sequencing platforms are unable to read the content of long strings that denote protein/gene sequences. One prototypical example of string reconstruction arises in DNA assembly: there, one creates multiple copies of the same long string and cuts the copies to read out short overlapping substrings that can be put together by matching their prefixes and suffixes. The reconstructed string may not be a perfect replica of the original string due to errors in the fragmentation and matching processes. Furthermore, for many strings unique reconstruction is inherently impossible. This represents a major issue for next generation sequencing technologies used in fundamental biological research, since in this setting it is impossible to ensure unambiguous results. The successful code designs pursued in this project can resolve reliability and content retrieval issues impeding implementations of emerging molecular computing and storage paradigms. This project is concerned with developing novel coding methods for unique reconstruction of strings or pools of strings based on their constituent substrings, subsequences and substring compositions. The techniques employed represent a combination of new graph-theoretic, combinatorial optimization and information theory approaches. In particular, the project will investigate the use of balanced partial de Bruijn strings for substring-based reconstruction, Catalan-like paths for multiset composition reconstruction as well as coded multi-trace reconstruction methods involving specialized modifications of deletion-correcting codes and superposition codes. The coding schemes will be tested on DNA-based and synthetic polymer-based data storage platforms under the development at the University of Illinois.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
The Gapped k-Deck Problem
k-Deck 缺口问题
DOI: --
发表时间: 2022
期刊: International Symposium on Information Theory and its Applications
影响因子: --
作者: [Rebecca Golm, Mina Nahvi]
通讯作者: Rebecca Golm, Mina Nahvi
Collaborative Research: CIF-Medium: Privacy-preserving Machine Learning on Graphs
Collaborative Research: CIF: Medium: Group testing for Real-Time Polymerase Chain Reactions: From Primer Selection to Amplification Curve Analysis
Collaborative Research: CIF: Medium: New Methods for Learning on Hypergraphs for Single-Cell Chromatin Data Analysis
CIF: Small: Collaborative Research:Leveraging Data Popularity in Distributed Storage Systems via Constrained Design Theory
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)