Geometric structures guided learning model and algorithms for bulk RNAseq data analysis
Geometric structures guided learning model and algorithms for bulk RNAseq data analysis
批准号:
10710214
负责人:
Duan Chen
金额:
$18.8万
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-28 至 2025-07-31
关键词:
AddressAlgorithmsAlzheimer&aposs DiseaseBenchmarkingBiologicalCell ExtractsCellsCharacteristicsComplexComputational algorithmComputer HardwareControl GroupsDataData AnalysesData CompressionData SetDiseaseExhibitsGene ExpressionGenesGeometryGraphHumanIndividualLaplacianLearningMathematicsMethodsModelingModernizationNoiseOutcomePharmaceutical PreparationsProcessRandomizedResearchSortingStructureTechniquesTissue SampleTissuesTranscription AlterationValidationWorkbiomarker identificationcell typecomputerized toolscostdata spacedifferential expressionfundamental researchgeometric structureinnovationlarge datasetsmathematical modeltranscriptome sequencing
中文摘要
许多疾病的潜在药物和治疗方法的发现在很大程度上取决于鉴别
在单个细胞类型内的疾病条件下表达的(DE)基因。虽然有可能
通过实验筛选出用于DE分析的单个细胞类型的细胞,通过计算利用大块组织
数据具有更高的可用性、更低的成本和更少的人工处理。关键的一步
这项研究的目的是(完全)解开特定细胞类型中的基因表达
异质的块状组织。完全反卷积可以看作是一种非负矩阵分解
(NMF)问题,然而,NMF是强不适定的,其不可分的解给出了巨大的挑战
在数据可解释性方面。这些挑战在不同的应用中是不同的,所以如果不采取特殊处理,
基因表达数据的完全去卷积结果将使DE分析几乎准确
不可能。在这项提议中,数学模型和相关的计算算法将是
为散体组织RNAseq分析的基础研究而建立的,为了更好的数据可解释性,
可靠性和效率。为了应对这一挑战,给定的批量组织数据集的几何结构
将首先进行探索,以确定组成细胞类型的标记基因。然后,建立了该模型
通过(1)强制NMF的弱可解性条件(由于噪声)和(2)执行几何
对已知数据空间的约束。这项工作的动机是许多
生物数据,在这些数据中,样本组织的表达水平显示出某些
基因。对于海量的生物数据,将发展随机快速计算算法。
在经过验证和基准测试后,该模型将应用于各种数据集的DE分析。
这一新提出的模型对于破译许多疾病中的细胞转录变化很重要。在……里面
建模策略,本研究提供了观察拓扑/几何的新视角
数据结构,实施相应的约束以增强问题的可解性和数据
可解释性。在计算上,本研究发展了非线性图的拉普拉斯正则化优化
结合随机压缩算法,能够以较低的存储空间处理海量数据。
要求高、复杂度低、适应现代计算机硬件结构。
AS
英文摘要
Discovering potential drugs and treatments of many diseases heavily depends on identifying differentially
expressed (DE) genes in disease conditions within individual cell types. While it is possible to
experimentally sort out cells of individual cell types for DE analysis, computationally leveraging bulk tissue
data has the advantage of greater availability, lower expenses, and less human handling. A critical step
toward this research is to (completely) deconvolute gene expressions in specific cell types from the
heterogeneous bulk tissues. Complete deconvolution can be viewed as a nonnegative matrix factorization
(NMF) problem, however, NMF is strongly ill-posed, and its non-separable solutions give great challenges
in data interpretability. These challenges vary in different applications, so if no special treatment is taken,
results from complete deconvolution of gene expression data will make accurate DE analysis almost
impossible. In this proposal, a mathematical model and associated computational algorithms will be
established for the fundamental research of bulk tissue RNAseq analysis, for better data interpretability,
reliability, and efficiency. To tackle this challenge, the geometric structure of the given bulk tissue data set
will be explored first to identify marker genes for the constituent cell types. Then the model is established
by (1) enforcing the weak solvability condition (because of noises) of NMF and (2) performing geometrical
constraints on the data space of knowns. This work is motivated by the common characteristics of many
biological data, in which expression levels across sample tissues exhibit strong correlations among certain
genes. For massive amount of biological data, stochastic fast computational algorithms will be developed.
After validation and benchmarking, the proposed model will be applied to DE analysis for various datasets.
This proposed new model is important to decipher cellular transcriptional alterations in many diseases. In
modeling strategies, this research provides a new perspective of observing topological/geometric
structures of data, enforcing the corresponding constraints to enhance problem solvability and data
interpretability. In computation, this research develops nonlinear graph Laplacian regularized optimization
associated with stochastic compression algorithms, which can process massive data with low storage.
requirement, low complexity, and adapt to modern structure of computer hardware.
As
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.3934/fods.2022013
发表时间:
2022-02
期刊:
Foundations of data science
影响因子:
2.3
作者:
[Duan Chen;Shaoyu Li;Xue Wang]
通讯作者:
Duan Chen;Shaoyu Li;Xue Wang
DOI:
10.1016/j.jcp.2023.112491
发表时间:
2023
期刊:
Journal of computational physics
影响因子:
4.1
作者:
[Chen,Duan]
通讯作者:
Chen,Duan
Geometric structures guided learning model and algorithms for bulk RNAseq data analysis
-
批准号:10592460
-
项目类别:
-
资助金额:$21.47万
-
财政年份:2022
-
负责人:Duan Chen
-
依托单位:
海外基金