The ENCODE Imputation Challenge: A critical assessment of methods for cross-cell type imputation of epigenomic profiles

The ENCODE Imputation Challenge: A critical assessment of methods for cross-cell type imputation of epigenomic profiles
复制标题

DOI:
10.1101/2022.07.30.502157
复制
发表时间:
2022-08
期刊:
bioRxiv
影响因子:
--
通讯作者:
Jacob Schreiber;C. Boix;Jin-Wook Lee;Hongyang Li;Yuanfang Guan;Chun-Chieh Chang;Jen-Chien Chang;Alex Hawkins-Hooker;Bernhard Schölkopf;Gabriele Schweikert;Mateo Rojas Carulla;Arif Canakoglu;Francesco Guzzo;Luca Nanni;M. Masseroli;Mark James Carman;Pietro Pinoli;Chenyang Hong;Kevin Y. Yip;J. P. Spence;S. S. Batra-S.;Yun S. Song;Shaun Mahony;Zheng Zhang;Wuwei Tan;Yang Shen;Yuanfei Sun;Minyi Shi;Jessika Adrian;R. Sandstrom;Nina P. Farrell;J. Halow;Kristen Lee;Lixia Jiang;Xinqiong Yang;Charles Epstein;J. Strattan;Michael Snyder;M. Kellis;W. S. Noble;A. Kundaje
Jacob Schreiber;C. Boix;Jin-Wook Lee;Hongyang Li;Yuanfang Guan;Chun-Chieh Chang;Jen-Chien Chang;Alex Hawkins-Hooker;Bernhard Schölkopf;Gabriele Schweikert;Mateo Rojas Carulla;Arif Canakoglu;Francesco Guzzo;Luca Nanni;M. Masseroli;Mark James Carman;Pietro Pinoli;Chenyang Hong;Kevin Y. Yip;J. P. Spence;S. S. Batra-S.;Yun S. Song;Shaun Mahony;Zheng Zhang;Wuwei Tan;Yang Shen;Yuanfei Sun;Minyi Shi;Jessika Adrian;R. Sandstrom;Nina P. Farrell;J. Halow;Kristen Lee;Lixia Jiang;Xinqiong Yang;Charles Epstein;J. Strattan;Michael Snyder;M. Kellis;W. S. Noble;A. Kundaje
中科院分区:
其他
文献类型:
--
作者:
Jacob Schreiber;C. Boix;Jin-Wook Lee;Hongyang Li;Yuanfang Guan;Chun-Chieh Chang;Jen-Chien Chang;Alex Hawkins-Hooker;Bernhard Schölkopf;Gabriele Schweikert;Mateo Rojas Carulla;Arif Canakoglu;Francesco Guzzo;Luca Nanni;M. Masseroli;Mark James Carman;Pietro Pinoli;Chenyang Hong;Kevin Y. Yip;J. P. Spence;S. S. Batra-S.;Yun S. Song;Shaun Mahony;Zheng Zhang;Wuwei Tan;Yang Shen;Yuanfei Sun;Minyi Shi;Jessika Adrian;R. Sandstrom;Nina P. Farrell;J. Halow;Kristen Lee;Lixia Jiang;Xinqiong Yang;Charles Epstein;J. Strattan;Michael Snyder;M. Kellis;W. S. Noble;A. Kundaje

文献摘要

相似文献

功能基因组学实验对于理解基因调控机制是非常宝贵的。然而,全面地进行所有这些实验,即使是在固定的样品和测定类型的集合中,在实践中通常是不可行的。彻底执行实验的一个有希望的替代方案是执行一组核心实验,然后使用机器学习方法来估算剩余的实验。然而,问题仍然是估算的质量,执行估算的最佳方法,甚至是什么样的性能指标有意义地评估这些模型的性能。在这项工作中,我们解决这些问题,通过全面分析23个插补模型提交给ENCODE插补挑战。我们发现,测量插补的质量比文献中报道的更具挑战性,并且受到三个因素的混淆:由于随着时间的推移数据收集和处理的差异而产生的主要分布变化,每个细胞类型的可用数据量以及性能指标之间的冗余。我们的系统分析提出了几个步骤,这些步骤是必要的,但也很简单,公平地评估这些模型的性能,以及在这一领域进行更强大的研究有前途的方向。
Functional genomics experiments are invaluable for understanding mechanisms of gene regulation. However, comprehensively performing all such experiments, even across a fixed set of sample and assay types, is often infeasible in practice. A promising alternative to performing experiments exhaustively is to, instead, perform a core set of experiments and subsequently use machine learning methods to impute the remaining experiments. However, questions remain as to the quality of the imputations, the best approaches for performing imputations, and even what performance measures meaningfully evaluate performance of such models. In this work, we address these questions by comprehensively analyzing imputations from 23 imputation models submitted to the ENCODE Imputation Challenge. We find that measuring the quality of imputations is significantly more challenging than reported in the literature, and is confounded by three factors: major distributional shifts that arise because of differences in data collection and processing over time, the amount of available data per cell type, and redundancy among performance measures. Our systematic analyses suggest several steps that are necessary, but also simple, for fairly evaluating the performance of such models, as well as promising directions for more robust research in this area.