Efficient Methods for Dimensionality Reduction ofSingle-Cell RNA-Sequencing Data
Efficient Methods for Dimensionality Reduction ofSingle-Cell RNA-Sequencing Data
批准号:
10356883
负责人:
James Michael Garritano
金额:
$5.18万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-03-16 至 2023-03-15
关键词:
AddressAdoptedAlgorithmsBiologicalCellsCodeCollectionCommunitiesComputer HardwareComputing MethodologiesConsensusDataData AnalysesData SetDevelopmentDimensionsDiseaseEvaluationFellowshipGaussian modelGenesHourHumanLanguageLearningLibrariesMathematicsMeasuresMentorshipMethodsModelingModernizationNamesNoiseNormal Statistical DistributionPaperPhysiciansPhysiologyPopulationPrincipal Component AnalysisProcessPublishingRNARandomizedResearch PersonnelResolutionRunningScientistSpeedStatistical BiasStatistical MethodsSystematic BiasTechniquesTechnologyTimeTissuesTrainingVariantVisualizationbasedesigndimensional analysisdistributed dataexperienceexperimental studyhigh dimensionalityimprovedinsightlaptopnon-Gaussian modelparallelizationprofessorsingle cell analysissingle-cell RNA sequencingstatisticssupercomputertheoriestooltranscriptometranscriptome sequencing
中文摘要
项目综述:单细胞RNA测序数据降维的有效方法
单细胞RNA测序是一项革命性的技术,使人类生理学和
疾病。单细胞RNA测序实验产生的数据集如此之大,以至于它们不可能
使用传统统计方法进行分析或可视化,直到使用
名为“降维”的技术。几乎所有的单细胞rna测序分析都是从
使用一种称为主成分分析(PCA)的技术来实现降维。
然而,单细胞RNA测序带来了独特的挑战,使得PCA变得困难。首先,这些东西的大小
数据集如此之大,以至于计算PCA需要专门的硬件和数小时。快速算法可实现
近似主成分分析已被证明显著加快了这一过程,但并未在
单细胞-RNA测序社区,部分原因是还没有并行算法在R
计算机语言。其次,主成分分析要求研究人员决定数据集的最终期望大小。
选择太小的尺码会导致丢弃有价值的生物学见解,而选择太大的尺码会导致放弃有价值的生物学见解
增加噪音。然而,对于如何选择单细胞rna的最佳大小,还没有达成共识。
测序,有证据表明这个大小可能被系统性地低估了。最后,主成分分析不能
直接应用于在单细胞RNA测序中测量的计数数据,因此研究人员必须首先应用
对其进行归一化处理的技术。目前该领域的标准是应用对数变换-
然而,最近的几项研究表明,对数变换在单细胞RNA中产生统计偏差
测序。在这项研究中,专门针对单细胞RNA执行PCA的快速方法-
将开发测序数据:1A)严格衡量变化后果的框架
几个公开可用的单细胞RNA测序数据集最终结果的前处理参数
使主成分分析能够在单细胞RNA测序数据上进行实验。1b)超高速、并行化
随机主成分分析的实现允许研究人员使用标准笔记本电脑快速执行主成分分析
单细胞RNA测序数据。2)表演时严格选择最终尺寸的技术
单细胞RNA测序数据集的主成分分析。3)一种转化单细胞的方法
RNA测序数据,以使其适当分布,从而能够正确使用PCA,而无需
招致统计偏差的。该奖学金还包括一份详细的培训计划,其中包含有价值的学习内容。
申请者作为一名能应用高等医学方法的内科科学家的发展经验
量纲统计在解决生物医学问题中的作用。
英文摘要
Project Summary: Efficient Methods for Dimensionality Reduction of Single-Cell RNA-Sequencing Data
Single-cell RNA-sequencing is a revolutionary technology enabling discoveries in human physiology and
disease. The datasets generated from single-cell RNA-sequencing experiments are so large that they cannot be
analyzed or visualized using traditional statistical methods until the datasets have been shrunk using a
technique named “dimensionality reduction.” Almost every analysis of single-cell RNA-sequencing begins
using a technique named principal component analysis (PCA) to accomplish dimensionality reduction.
However, single-cell RNA-sequencing presents unique challenges making PCA difficult. First, the size of these
datasets is so large that computing PCA requires specialized hardware and multiple hours. Fast algorithms to
approximate PCA have been shown to dramatically speed up this process, but have not proliferated in the
single cell-RNA sequencing community, in part because no parallelized algorithm has been written in the R
computing language. Second, PCA requires the researcher to decide the final desired size of the dataset.
Choosing too small of a size results in discarding valuable biological insights, while choosing too large a size
increases the noise. However, there is no consensus on how to pick the optimal size for single-cell RNA
sequencing, and there is evidence that this size might be systematically underestimated. Lastly, PCA cannot be
applied directly to the count-data measured in single cell RNA sequencing, so researchers must first apply a
preprocessing technique to normalize it. The current standard in the field is to apply the log transform –
however, several recent studies have shown that the log transform creates statistical biases in single-cell RNA
sequencing. In this fellowship, specifically tailored, fast methods for performing PCA on single-cell RNA-
sequencing data will be developed: 1a) A framework to rigorously measure the consequence of changing
preprocessing parameters on the final results of several publicly available single cell RNA sequencing datasets
to enable experimentation of PCA on single-cell RNA-sequencing data. 1b) An ultra-fast, parallelized
implementation of randomized PCA allowing researchers using standard laptops to rapidly perform PCA on
single cell RNA sequencing data. 2) A technique for rigorously choosing the final size when performing
principal component analysis for single-cell RNA-sequencing datasets. 3) A method for transforming single-cell
RNA-sequencing data so that it becomes appropriately distributed enabling proper usage of PCA without
incurring statistical biases. This fellowship also includes a detailed training plan with valuable learning
experiences for the applicant’s development as a physician-scientist who can apply methods from high
dimensional-statistics to solving biomedical problems.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1002/cncy.22399
发表时间:
2021-05
期刊:
Cancer cytopathology
影响因子:
3.4
作者:
[Abi-Raad R, Prasad ML, Gilani S, Garritano J, Barlow D, Cai G, Adeniran AJ]
通讯作者:
Adeniran AJ
DOI:
10.1002/cncy.22537
发表时间:
2022-04
期刊:
CANCER CYTOPATHOLOGY
影响因子:
3.4
作者:
[Gilani, Syed M., Abi-Raad, Rita, Garritano, James, Cai, Guoping, Prasad, Manju L., Adeniran, Adebowale J.]
通讯作者:
Adeniran, Adebowale J.
Anaplastic Thyroid Carcinoma: Cytomorphologic Features on Fine-Needle Aspiration and Associated Diagnostic Challenges.
甲状腺未分化癌:细针抽吸的细胞形态学特征及相关诊断挑战。
DOI:
10.1093/ajcp/aqab159
发表时间:
2022
期刊:
American journal of clinical pathology
影响因子:
3.5
作者:
[Podany,Peter, Abi-Raad,Rita, Barbieri,Andrea, Garritano,James, Prasad,ManjuL, Cai,Guoping, Adeniran,AdebowaleJ, Gilani,SyedM]
通讯作者:
Gilani,SyedM
海外基金