课题基金 / 基金详情

CAREER: Algorithms and Tools for Allele-Specific Transcript Assembly

CAREER: Algorithms and Tools for Allele-Specific Transcript Assembly
职业:等位基因特异性转录本组装的算法和工具
批准号:
2145171
负责人:
Mingfu Shao
金额:
$74.99万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-01 至 2027-06-30

项目摘要

项目成果

Mingfu Shao的其他基金

相似基金

相关文献

中文摘要
翻译
该奖项的全部或部分资金来自《2021年美国救援计划法案》(公法117-2)。许多生物,包括人类,都有两组染色体,一组来自母亲,另一组来自父亲,它们以同源对的形式存在。通常观察到,在两条同源染色体的同一位置上的两个不同的基因(或等位基因)可能产生不平衡的基因产物(即mRNAs)。这种现象被称为等位基因特异性表达(ASE)。已知ASE与多种表型密切相关,并可促进癌症易感性。ASES提供了一个重要的生物标记物来源,可能被用作表型生物标记物或用于疾病诊断。此外,ASE分析是确定表达数量性状基因座(EQTL)和研究各种生物学过程的强大分析工具,如印迹、蛋白质截断变体和X染色体失活。最近建立的RNA测序技术(RNA-seq)为定量测定ASE提供了一种准确而有效的方法。然而,由当前技术产生的测序读数并不是全长的。因此,需要计算方法来重建存在于同源染色体上的两个不同等位基因表达的全长mRNAs,这一问题被称为等位基因特异性转录组装。等位基因特异性转录组装是非常困难的。等位基因特异性转录本的组装是困难的,因为它需要同时处理突变和剪接连接,同时推断未知数量的全长转录本及其丰度。该项目旨在开发适用于短读、长读和单细胞RNA序列数据的准确的等位基因特异性转录本组装方法。具体地说,这些研究人员首先研究如何使用阶段性SNPs(即突变)来改善等位基因特异性组装。它们表明,在所谓的变异拼接图中,相对应的SNP可以等价地表示为不相容的顶点对。然后提出了求解包含不相容对的公式的启发式算法。长程信息将被用于配对/多端RNA-SEQ数据中,以改进等位基因特异性组装,并使用一种算法将可变剪接图分解为路径,同时完全保留配对/多端约束。等位基因特异性组装将在存在结构变异的情况下得到解决,这是研究癌症的关键情景。一种新的数据结构将对SNP、选择性剪接和结构变异全部建模。还提出了一种新的算法来识别等位基因特异的结构变异,这是独立的兴趣,但也导致了等位基因特异组装的两步算法。协变量自适应多重假设检验将控制假阳性率。实施这些算法可以为各种类型的数据生成准确的等位基因特定转录本汇编器,并为更广泛的使用提供新的ASE分析流水线。建议的研究与教育活动很好地结合在一起。高中课程将侧重于使用图表结构--数学和计算机科学中的一个关键抽象--来对生物数据进行建模。高中教师将有机会进行与成绩单汇编相关的跨学科研究。将开发新的本科课程,重点是提高学生建模和解决现实世界问题的能力。还将努力吸引来自代表性不足群体的本科生和研究生参与研究机会。该项目的结果可以在PI的网站上找到:https://sites.psu.edu/mxs2589.This奖反映了美国国家科学基金会的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This award is funded in whole or in part under the American Rescue Plan Act of 2021 (Public Law 117-2).Many organisms, including humans, have two sets of chromosomes, one from the mother and one from the father, that exist as homologous pairs. It is commonly observed that, the two distinct genes (or alleles) at the same location of two homologous chromosomes, may produce imbalanced gene products (i.e., mRNAs). This phenomenon is called allele-specific expression (ASE). ASE has been known to be closely related to multiple phenotypes and can contribute to cancer susceptibility. ASEs offer an important source of biomarkers that could be potentially used as phenotypic biomarkers or for disease diagnosis. Additionally, ASE analysis serves as a powerful analytical tool to determine expression quantitative trait locus (eQTL) and to study a variety of biological processes such as imprinting, protein-truncating variants, and X-chromosome inactivation. The recently established RNA-sequencing technology (RNA-seq) provides an accurate and efficient way to quantitatively measure ASE. However, the sequencing reads generated from current technologies are not full-length. Hence, computational methods are needed to reconstruct the full-length mRNAs expressed from the two different alleles that exist on homologous chromosomes, a problem referred to as allele-specific transcript assembly. Allele-specific transcript assembly is exceedingly difficult. Allel-specific transcript assembly is difficult because it requires simultaneously threading mutations and splice junctions while inferring unknown number of full-length transcripts and their abundances. This project aims to develop accurate allele-specific transcript assembly methods that are applicable to short-reads, long-reads, and single-cell RNA-seq data. Specifically, these investigators first tackle how to use phased SNPs (i.e., mutations) to improve allele-specific assembly. they show that, phased SNPs can be equivalently represented as incompatible pairs of vertices in a so-called variant splice graph. Then heuristics are proposed to solve a formulation with incompatible pairs included. Long-range information will be used in paired-/multi-end RNA-seq data to improve allele-specific assembly, with an algorithm that decomposes the variant splice graph into paths while fully preserving the paired-/multi-end constraints. The allele-specific assembly will be solved in the presence of structure variations, a crucial scenario in studying cancer. A new data structure will model SNPs, alternative splicing, and structure variations all together. A new algorithms is also proposed to identify allele-specific structure variations, which is of independent interests but also leads to a two-step algorithm for allele-specific assembly. Covariate-adaptive multiple hypothesis testing will control false positive rates. Implementing these algorithms result in accurate allele-specific transcript assemblers for various types of data and a new ASE analysis pipeline for broader use. The proposed research is well integrated with educational activities. High-school curricula will be developed that focus on using graph structure--a key abstract in mathematics and computer science--to model biological data. High school teachers will be provided opportunities to conduct interdisciplinary research related to transcript assembly. New undergraduate course will be developed with a focus to enhance students’ ability in modeling and solving real-world problems. Efforts will also be made to engage undergraduates and graduate students from underrepresented groups in research opportunities. The results of the project can be found at the PI’s website: https://sites.psu.edu/mxs2589.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
BBSRC-NSF/BIO: IIBR Informatics: Collaborative Research: Inference of isoform-level regulatory infrastructures with studies in steroid-producing cells
海外基金