Visualizing the structure of RNA-seq expression data using grade of membership models

Visualizing the structure of RNA-seq expression data using grade of membership models
复制标题

DOI:
10.1371/journal.pgen.1006599
复制
发表时间:
2017-03-01
期刊:
影响因子:
4.5
通讯作者:
Stephens, Matthew
Stephens, Matthew
中科院分区:
生物学2区
文献类型:
--
作者:
Dey, Kushal K.;Hsiao, Chiaowen Joyce;Stephens, Matthew

文献摘要

被引文献

相似文献

隶属度模型(Grade of Membership Models),也被称为“混合模型”、“主题模型”或“潜在狄利克雷分配”,是聚类模型的一种推广,它允许每个样本在多个聚类中具有成员资格。这些模型被广泛用于群体遗传学,以模拟具有来自多个“群体”的祖先的混合个体,并在自然语言处理中对具有来自多个“主题”的单词的文档进行建模。在这里,我们说明了这些模型聚类RNA-seq基因表达数据样本的潜力,这些数据是在批量样本或单细胞上测量的。我们还提供了方法来帮助解释集群,通过识别在每个集群中独特表达的基因。通过将这些方法应用于几个示例RNA-seq应用程序,我们证明了它们在识别和总结结构和异质性方面的实用性。应用于GTEx项目53个人体组织的数据,该方法突出了生物相关组织之间的相似性,并确定了重现已知生物学的独特表达基因。应用于小鼠植入前胚胎的单细胞表达数据,该方法突出了通过早期胚胎发育阶段的离散和连续变化,并突出了从生殖细胞发育,通过压实和桑椹胚形成,到胚泡阶段内细胞团和滋养层形成的各种相关过程D中涉及的基因。这些方法在Bioconductor包CountClust中实现。
Grade of membership models, also known as '' admixture models '', '' topic models '' or '' Latent Dirichlet Allocation '', are a generalization of cluster models that allow each sample to have membership in multiple clusters. These models are widely used in population genetics to model admixed individuals who have ancestry from multiple '' populations '', and in natural language processing to model documents having words from multiple '' topics ''. Here we illustrate the potential for these models to cluster samples of RNA-seq gene expression data, measured on either bulk samples or single cells. We also provide methods to help interpret the clusters, by identifying genes that are distinctively expressed in each cluster. By applying these methods to several example RNA-seq applications we demonstrate their utility in identifying and summarizing structure and heterogeneity. Applied to data from the GTEx project on 53 human tissues, the approach highlights similarities among biologicallyrelated tissues and identifies distinctively-expressed genes that recapitulate known biology. Applied to single-cell expression data from mouse preimplantation embryos, the approach highlights both discrete and continuous variation through early embryonic development stages, and highlights genes involved in a variety of relevant processes D from germ cell development, through compaction and morula formation, to the formation of inner cell mass and trophoblast at the blastocyst stage. The methods are implemented in the Bioconductor package CountClust.