MultiPLIER: A Transfer Learning Framework for Transcriptomics Reveals Systemic Features of Rare Disease

MultiPLIER: A Transfer Learning Framework for Transcriptomics Reveals Systemic Features of Rare Disease
复制标题

DOI:
10.1016/j.cels.2019.04.003
复制
发表时间:
2019-05-22
期刊:
影响因子:
9.3
通讯作者:
Greene, Casey S.
Greene, Casey S.
中科院分区:
生物学1区
文献类型:
--
作者:
Taroni, Jaclyn N.;Grayson, Peter C.;Greene, Casey S.

文献摘要

被引文献

相似文献

个体研究人员生成的大多数基因表达数据集太小,无法完全受益于无监督机器学习方法。在罕见病的情况下,即使将多项研究结合起来,可用的病例也可能太少。为了应对这一挑战,我们利用迁移学习来提取协调的表达模式,并使用学习到的模式来分析小型罕见疾病数据集。我们在一个包含多个实验、组织和生物条件的大型公共数据库上训练了一个路径级信息提取器(pathway-level information extractor,简称PIMER)模型,然后将该模型转移到小数据集上,我们称之为MultiPIMER。从公共数据汇编构建的模型包括与已知生物因素高度一致的特征,并且比从单个数据集或条件构建的模型更全面。当转移到罕见疾病数据集时,模型比仅在给定数据集上训练的模型更有效地描述与疾病严重程度相关的生物过程。
Most gene expression datasets generated by individual researchers are too small to fully benefit from unsupervised machine-learning methods. In the case of rare diseases, there may be too few cases available, even when multiple studies are combined. To address this challenge, we utilize transfer learning to extract coordinated expression patterns and use learned patterns to analyze small rare disease datasets. We trained a pathway-level information extractor (PLIER) model on a large public data compendium comprising multiple experiments, tissues, and biological conditions and then transferred the model to small datasets in an approach we call MultiPLIER. Models constructed from the public data compendium included features that aligned well to known biological factors and were more comprehensive than those constructed from individual datasets or conditions. When transferred to rare disease datasets, the models describe biological processes related to disease severity more effectively than models trained only on a given dataset.