Large-scale mapping of mammalian transcriptomes identifies conserved genes associated with different cell states.

Large-scale mapping of mammalian transcriptomes identifies conserved genes associated with different cell states.
复制标题

哺乳动物转录组的大规模作图鉴定了与不同细胞状态相关的保守基因

DOI:
10.1093/nar/gkw1256
复制
发表时间:
2017-02-28
影响因子:
14.9
通讯作者:
Li JJ
Li JJ
中科院分区:
生物学2区
文献类型:
--
作者:
Yang Y;Yang YT;Yuan J;Lu ZJ;Li JJ

文献摘要

相似文献

仅根据基因表达数据区分细胞状态仍然是一项具有挑战性的任务。这是真实的,甚至在一个物种内的分析。在跨物种比较中,不同小组获得的结果差异很大。在这里,我们整合了来自四种哺乳动物物种的40多种细胞和组织类型的RNA-seq数据,以确定作为每个物种特定细胞状态指标的相关基因集。我们采用统计方法,TROM,以确定蛋白质编码和非编码指标。接下来,我们使用这些指示基因绘制每个物种内以及物种之间的细胞状态。我们概括了相关细胞和组织类型之间已知的表型相似性,并揭示了它们相似性的分子基础。我们还报告了几种组织和细胞类型与功能支持之间的新关联。此外,我们发现的保守相关基因是研究细胞分化和重编程的良好资源。最后,长的非编码RNA可以很好地作为相关基因来指示细胞状态。根据这些非编码相关基因与蛋白质编码基因的共表达,我们进一步推测了它们的生物学功能。这项研究表明,将统计建模与公共RNA-seq数据相结合,可以有效地提高我们对细胞身份控制的理解。
Distinguishing cell states based only on gene expression data remains a challenging task. This is true even for analyses within a species. In cross-species comparisons, the results obtained by different groups have varied widely. Here, we integrate RNA-seq data from more than 40 cell and tissue types of four mammalian species to identify sets of associated genes as indicators for specific cell states in each species. We employ a statistical method, TROM, to identify both protein-coding and non-coding indicators. Next, we map the cell states within each species and also between species using these indicator genes. We recapitulate known phenotypic similarity between related cell and tissue types and reveal molecular basis for their similarity. We also report novel associations between several tissues and cell types with functional support. Moreover, our identified conserved associated genes are found to be a good resource for studying cell differentiation and reprogramming. Lastly, long non-coding RNAs can serve well as associated genes to indicate cell states. We further infer the biological functions of those non-coding associated genes based on their co-expressed protein-coding genes. This study demonstrates that combining statistical modeling with public RNA-seq data can be powerful for improving our understanding of cell identity control.