EFFICIENT METHODS FOR CALIBRATION, CLUSTERING, VISUALIZATION AND IMPUTATION OF LARGE scRNA-seq DATA
EFFICIENT METHODS FOR CALIBRATION, CLUSTERING, VISUALIZATION AND IMPUTATION OF LARGE scRNA-seq DATA
批准号:
10335252
负责人:
Yuval Kluger
金额:
$40.04万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2019
资助国家:
美国
项目状态:
未结题
起止时间:
2019-05-01 至 2025-01-31
关键词:
3-DimensionalAddressAdoptedAlgorithmsAttenuatedBenchmarkingBig Data MethodsBig Data to KnowledgeBiologicalBiologyCalibrationCellsComputational BiologyDNA MethylationDataData AnalysesData AnalyticsData SetDetectionDiffusionDimensionsDropoutEmerging TechnologiesExcisionFundingGenerationsGenesGenomic approachGraphImmuneLaplacianLearningMapsMeasurementMethodsModalityNeuronsNoisePathogenesisPhenotypePopulationProbabilityProceduresRecoveryResearchResearch PersonnelSamplingScienceSeriesSignal TransductionSpeedStructureSystemTechniquesUnited States National Institutes of HealthValidationVariantVisualizationanalytical methodartificial neural networkbasebiomarker discoverycell typecomputerized toolsdeep learningdeep neural networkdensityexperimental studyhematopoietic differentiationhuman diseaseimprovedinsightkernel methodslarge datasetslearning networkmultidimensional dataneural networknovelprototyperesponsesingle cell analysissingle-cell RNA sequencingtheoriestooltranscriptome sequencing
中文摘要
单细胞rna-seq(scrna-seq)图谱为进行详细的细胞学研究提供了前所未有的机会。
细胞亚群分析。实现scRNA-seq在生物医学研究和生物标记物中的前景
发现需要强大的计算方法来支持罕见和意外表型的检测
细胞反应。当前scRNA-seq的归属、校准、聚类和可视化方法
数据面临着诸如错误输入未表达的基因、线性限制等挑战
去除多变量批次效应的假设,以及聚类和降维的低效
非常大的数据集的方法。我们开发了谱、神经网络和快速多极子方法
(FMM)适合在scRNA-seq和其他高吞吐量的背景下解决这些问题的原型
数据背景,并提议进一步开发和调整这些方法以用于scRNA-seq数据分析。我们的团队
数据分析和计算生物学专家小组目前通过NIH BD2K计划获得资金,以
开发对生物医学科学具有广泛适用性的新型大数据工具和方法。这一努力
证明了神经网络、频谱和调和的极高效可扩展原型的可行性
适用于校准、降维和可视化高维数据的分析技术,
寻找本征状态概率密度,并共同组织细胞、标记和样本。我们建议
在包括基质在内的单细胞RNA-SEQ研究中使用的现有分析方法的实质性进展
结合矩阵补全和统计的稀疏和噪声scRNA-seq数据恢复方法
技术(目标1A),以及基于我们的无监督MMD-ResNet神经网络原型和
最优运输理论(目标1B)。我们将开发FMM方法的一个变体来加速计算
T分布随机邻近嵌入(t-SNE)可视化技术的排斥项,其
将改进我们目前最快的基于t-SNE FFT的Fit-SNE原型,并开发新的可靠的近似
加速t-SNE等聚类吸引项计算的最近邻方法
算法(目标2A)。我们将进一步开发t-SNE的其他变种,以实现更好的分离
在细胞亚群集群之间(后期夸大)和使用1D t-SNE进行更好的可视化
热图基因-细胞表达(目标2A)。我们将采用我们的高效神经网络方法SpectralNet,
用于计算大型数据集的图的拉普拉斯特征向量。这将使频谱计算成为可能
聚类、扩散图和流形学习在许多scRNA研究中使用,但目前
仅限于中等数量的单细胞(目标2B)。最后,我们将开发一种基于核的差分
描述生物条件差异的丰度算法(目标2C)。我们将领养
适当的抽样方法,以显著改进目前的方法。
英文摘要
Single cell RNA-seq (scRNA-seq) profiling provides an unprecedented opportunity to conduct detailed cellular
analysis of cell subpopulations. Fulfilling the promise of scRNA-seq for biomedical studies and biomarker
discovery requires robust computational approaches to support detection of rare phenotypes and unanticipated
cellular responses. Current approaches for imputation, calibration, clustering and visualizing of scRNA-seq
data suffer from challenges such as erroneous imputing of non-expressed genes, limitation of linear
assumptions in removal of multivariate batch effects, and inefficiencies of clustering and dimensional reduction
methods of very large datasets. We have developed spectral, neural network, and Fast Multipole Methods
(FMM) prototypes suitable for addressing these issues in the context of scRNA-seq and other high throughput
data contexts and propose to further develop and adapt these methods to scRNA-seq data analysis. Our team
of experts on data analytics and computational biology is currently funded through the NIH BD2K initiative to
develop novel big data tools and methods that have broad applicability to biomedical science. This effort
proved the feasibility of extremely efficient scalable prototypes of neural network, spectral, and harmonic
analysis techniques suitable for calibrating, reducing the dimensionality and visualizing high dimensional data,
finding intrinsic state-probability densities, and co-organizing cells, markers and samples. We propose
substantial advances over existing analytical procedures used in single cell RNA-seq studies including matrix
recovery approaches for the sparse and noisy scRNA-seq data by combining matrix completion and statistical
techniques (Aim 1A), and calibration based on our unsupervised MMD-ResNet neural network prototype and
optimal transport theory (Aim 1B). We will develop a variant of the FMM approach to speed up the calculation
of the repulsion term of the t-distributed stochastic neighbor embedding (t-SNE) visualization technique, which
will improve our current fastest t-SNE FFT-based FIt-SNE prototype, and develop new reliable approximate
nearest neighbors approaches to speed up the computation of the attraction term of t-SNE and other clustering
algorithms (Aim 2A). Our additional variants of t-SNE will be further developed to allow better separation
between clusters of cell subpopulations (late exaggeration) and better visualization using 1D t-SNE for
heatmap gene-cell representation (Aim 2A). We will adapt SpectralNet, our efficient neural network approach,
for computing graph Laplacian eigenvectors for large datasets. This will enable computation of spectral
clustering, diffusion maps and manifold learning that are utilized in many scRNA studies but are currently
limited to a moderate number of single cells (Aim 2B). Finally, we will develop a kernel based differential
abundance algorithm to characterize differences between biological conditions (Aim 2C). We will adopt
appropriate sampling approaches to significantly improve current methods.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1093/imaiai/iaz018
发表时间:
2017-09
期刊:
Information and inference : a journal of the IMA
影响因子:
--
作者:
[Xiuyuan Cheng;A. Cloninger;R. Coifman]
通讯作者:
Xiuyuan Cheng;A. Cloninger;R. Coifman
DOI:
10.1109/tit.2022.3175691
发表时间:
2019-09
期刊:
IEEE Transactions on Information Theory
影响因子:
2.5
作者:
[Xiuyuan Cheng;A. Cloninger]
通讯作者:
Xiuyuan Cheng;A. Cloninger
DOI:
10.3390/cancers12092551
发表时间:
2020-09-08
期刊:
Cancers
影响因子:
5.2
作者:
[Marczyk M, Patwardhan GA, Zhao J, Qu R, Li X, Wali VB, Gupta AK, Pillai MM, Kluger Y, Yan Q, Hatzis C, Pusztai L, Gunasekharan V]
通讯作者:
Gunasekharan V
Core C
-
批准号:10553035
-
项目类别:
-
资助金额:$45.75万
-
财政年份:2022
-
负责人:Yuval Kluger
-
依托单位:
Core C
-
批准号:10675116
-
项目类别:
-
资助金额:$48.65万
-
财政年份:2022
-
负责人:Yuval Kluger
-
依托单位:
Core D: Data Analysis Core
-
批准号:10384402
-
项目类别:
-
资助金额:$36.01万
-
财政年份:2021
-
负责人:Yuval Kluger
-
依托单位:
Core D: Data Analysis Core
-
批准号:10689282
-
项目类别:
-
资助金额:$33.51万
-
财政年份:2021
-
负责人:Yuval Kluger
-
依托单位:
EFFICIENT METHODS FOR CALIBRATION, CLUSTERING, VISUALIZATION AND IMPUTATION OF LARGE scRNA-seq DATA
-
批准号:9920743
-
项目类别:
-
资助金额:$39.11万
-
财政年份:2019
-
负责人:Yuval Kluger
-
依托单位:
EFFICIENT METHODS FOR CALIBRATION, CLUSTERING, VISUALIZATION AND IMPUTATION OF LARGE scRNA-seq DATA
-
批准号:9764594
-
项目类别:
-
资助金额:$41.75万
-
财政年份:2019
-
负责人:Yuval Kluger
-
依托单位:
EFFICIENT SPECTRAL APPROACHES FOR FINDING UNDERLYING STRUCTURES IN BIG DATA
-
批准号:9278252
-
项目类别:
-
资助金额:$41.0万
-
财政年份:2016
-
负责人:Yuval Kluger
-
依托单位:
Co-ordination of recombination and allelic exclusion at IgH and Igk loci
-
批准号:8740626
-
项目类别:
-
资助金额:$8.48万
-
财政年份:2014
-
负责人:Yuval Kluger
-
依托单位:
Role of ATM and RAG in maintaining genome stability during Tcra/d rearrangement.
-
批准号:8707743
-
项目类别:
-
资助金额:$39.25万
-
财政年份:2013
-
负责人:Yuval Kluger
-
依托单位:
Role of ATM and RAG in maintaining genome stability during Tcra/d rearrangement.
-
批准号:8513573
-
项目类别:
-
资助金额:$42.25万
-
财政年份:2012
-
负责人:Yuval Kluger
-
依托单位:
Role Of Nuclear Organization In Protecting Genome Stability During Recombination
-
批准号:8505572
-
项目类别:
-
资助金额:$40.48万
-
财政年份:2008
-
负责人:Yuval Kluger
-
依托单位:
Role Of Nuclear Organization In Protecting Genome Stability During Recombination
-
批准号:8703127
-
项目类别:
-
资助金额:$38.74万
-
财政年份:2008
-
负责人:Yuval Kluger
-
依托单位:
Role Of Nuclear Organization In Protecting Genome Stability During Recombination
-
批准号:9040987
-
项目类别:
-
资助金额:$38.74万
-
财政年份:2008
-
负责人:Yuval Kluger
-
依托单位:
海外基金