课题基金 / 基金详情

项目摘要

项目成果

Gen Li的其他基金

相似基金

相关文献

中文摘要
翻译
项目总结 RNA测序(RNA-Seq)分析为了解基因功能提供了重要手段。高吞吐量 RNA-Seq数据经常在多个条件下从同一组样本中测量。例如, 在NIH共同基金的基因类型-组织表达(GTEx)项目中,来自不同组织的样本 从每个死后捐赠者身上收集来进行测序。关于紫外线(UV)辐射的另一项研究,皮肤 来自同一组受试者的角质形成细胞之前暴露于不同的辐射剂量和持续时间 测序。这种共同样本、多条件RNA-Seq数据具有在 样本和条件,并有可能提供对基因功能的关键见解。然而,尽管 收集这些数据的努力很大,但缺乏分析方法和计算工具来最大限度地 他们的潜力。重要任务,如缺失数据补充、功能基因模块识别和 关联分析仍然没有得到解决。在这项提议中,我们将建立一个创新和强大的范式 分析多条件RNA-Seq数据,从而提高我们对基因功能的理解。以杠杆作用 同时跨条件、样本和基因的信息,我们建议将RNA-Seq数据建模为 多向张量阵列。我们将开发适合于读取计数的新的张量方法和理论 数据。特别地,我们的第一个目标是扩展针对分块丢失的RNA-Seq数据的张量补全方法 推卸责任。通过将未观察到的样本建模为张量中缺失的块,我们将聚合信息 沿着不同的模式(受试者、条件、基因)来归因于缺失值。第二个目标是发展 灵活的张量共聚类方法,它同时对基因、样本和条件进行聚类,用于联合 表达基因模块鉴定。第三个目标是建立新的张量响应回归模型 将基因模块与基因和协变量相关联,这将为研究基因调控提供见解 AS表达数量性状基因座(EQTL)。最后,在第四个目标中,我们将开发可伸缩的统计 软件来实施所提出的方法,并使其更广泛地适用。我们将应用 方法对GTEx多组织数据和UV多状态数据进行处理,获得对基因的新见解 表达和调控。这项拟议的研究可能会改变我们分析多条件RNA的方式- SEQ数据,并加强我们对人类基因组学及其与公共健康的关系的理解。
英文摘要
PROJECT SUMMARY RNA-Sequencing (RNA-Seq) analysis provides a critical means to understand gene functions. High-throughput RNA-Seq data are frequently measured under multiple conditions from the same set of samples. For example, in the NIH Common Fund’s Genotype-Tissue Expression (GTEx) project, samples from different tissues are collected from each post-mortem donor for sequencing. For another study on ultraviolet (UV) radiation, skin keratinocytes from the same set of subjects are exposed to different radiation doses and durations before sequencing. Such common-sample, multi-condition RNA-Seq data have information shared across both samples and conditions, and have the potential to provide key insights into gene functions. However, despite great endeavors to collect such data, there is a lack of analytical methods and computational tools to maximize their potential. Important tasks such as missing data imputation, functional gene module identification and association analysis remain unaddressed. In this proposal, we will build an innovative and powerful paradigm to analyze multi-condition RNA-Seq data and thus improve our understanding of gene functions. To leverage information across conditions, samples and genes simultaneously, we propose to model RNA-Seq data as multi-way tensor arrays. We will develop novel tensor methods and theory that are appropriate for read count data. In particular, our first aim is to extend tensor completion methods for block-wise missing RNA-Seq data imputation. By modeling unobserved samples as missing blocks in a tensor, we will aggregate information along different modes (subjects, conditions, genes) to impute missing values. The second aim develops flexible tensor co-clustering methods, which simultaneously cluster genes, samples and conditions, for co- expressed gene module identification. The third aim is to build new tensor response regression models to associate gene modules with genotype and covariates which will provide insights into genetic regulation such as expression quantitative trait loci (eQTL). Finally, in the fourth aim, we will develop scalable statistical software to implement the proposed methods and make them more broadly applicable. We will apply the methods to the GTEx multi-tissue data and UV multi-condition data, and gain novel insights into gene expression and regulation. The proposed research will likely transform how we analyze multi-condition RNA- Seq data and enhance our understanding of human genomics and its relation to public health.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Tensor Array Methods for RNA-Seq Analysis
Tensor Array Methods for RNA-Seq Analysis
Tensor Array Methods for RNA-Seq Analysis
海外基金