Representing high throughput expression profiles via perturbation barcodes reveals compound targets.

Representing high throughput expression profiles via perturbation barcodes reveals compound targets.
复制标题

DOI:
10.1371/journal.pcbi.1005335
复制
发表时间:
2017-02
影响因子:
4.3
通讯作者:
Tudor M
Tudor M
中科院分区:
生物学2区
文献类型:
--
作者:
Filzen TM;Kutchukian PS;Hermes JD;Li J;Tudor M

文献摘要

被引文献

相似文献

高通量mRNA表达谱可用于表征细胞培养模型对扰动(例如药理学调节剂和遗传扰动)的响应。随着分析活动范围的扩大,重要的是要以捕获重要生物信号的方式对所得数据进行均质化、总结和分析,而不管各种噪声源,例如批次效应和随机变化。我们使用L1000平台对数千种化合物治疗中的978个代表性基因进行了大规模分析。在这里,描述了一种使用深度学习技术将地标基因的表达变化转换为扰动条形码的方法,该扰动条形码揭示了基础数据的重要特征,在揭示重要的生物学见解方面比原始数据表现得更好。条形码捕获化合物结构和靶标信息,并预测化合物的高通量筛选混杂性,达到比原始数据测量更高的程度,表明该方法揭示了表达数据的潜在因素,否则这些因素会被噪声纠缠或掩盖。此外,我们证明,来自扰动条形码的可视化可以用来更灵敏地分配功能未知的化合物通过内疚的协会的方法,我们用它来预测和实验验证的活性的化合物的MAPK途径。深度度量学习在大规模化学遗传学项目中的应用突出了这种方法和相关方法在从大数据(有时是噪声数据)中提取见解和可测试假设方面的实用性。小分子或生物制剂的作用可以通过它们对细胞基因表达谱的影响来测量。几十年来,这种实验一直是用小的、集中的样本集进行的。技术的进步现在允许这种方法在每年数万个样本的规模上使用。随着数据集大小的增加,由于实验和生物噪声以及表型不明显的事实,它们的分析变得更加困难。我们证明,使用为深度学习开发的工具,可以生成用于表达实验的“条形码”,这些实验可用于简单,高效和可重复地将细胞处理的表型效应表示为100个1和0的字符串。我们发现这种条形码在捕获潜在生物学方面比原始基因表达水平做得更好,并继续表明它可用于识别未表征分子的靶标。
High throughput mRNA expression profiling can be used to characterize the response of cell culture models to perturbations such as pharmacologic modulators and genetic perturbations. As profiling campaigns expand in scope, it is important to homogenize, summarize, and analyze the resulting data in a manner that captures significant biological signals in spite of various noise sources such as batch effects and stochastic variation. We used the L1000 platform for large-scale profiling of 978 representative genes across thousands of compound treatments. Here, a method is described that uses deep learning techniques to convert the expression changes of the landmark genes into a perturbation barcode that reveals important features of the underlying data, performing better than the raw data in revealing important biological insights. The barcode captures compound structure and target information, and predicts a compound’s high throughput screening promiscuity, to a higher degree than the original data measurements, indicating that the approach uncovers underlying factors of the expression data that are otherwise entangled or masked by noise. Furthermore, we demonstrate that visualizations derived from the perturbation barcode can be used to more sensitively assign functions to unknown compounds through a guilt-by-association approach, which we use to predict and experimentally validate the activity of compounds on the MAPK pathway. The demonstrated application of deep metric learning to large-scale chemical genetics projects highlights the utility of this and related approaches to the extraction of insights and testable hypotheses from big, sometimes noisy data. The effects of small molecules or biologics can be measured via their effect on cells’ gene expression profiles. Such experiments have been performed with small, focused sample sets for decades. Technological advances now permit this approach to be used on the scale of tens of thousands of samples per year. As datasets increase in size, their analysis becomes qualitatively more difficult due to experimental and biological noise and the fact that phenotypes are not distinct. We demonstrate that using tools developed for deep learning it is possible to generate ‘barcodes’ for expression experiments that can be used to simply, efficiently, and reproducibly represent the phenotypic effects of cell treatments as a string of 100 ones and zeroes. We find that this barcode does a better job of capturing the underlying biology than the original gene expression levels, and go on to show that it can be used to identify the targets of uncharacterized molecules.