Weighted elastic net for unsupervised domain adaptation with application to age prediction from DNA methylation data

Weighted elastic net for unsupervised domain adaptation with application to age prediction from DNA methylation data
复制标题

DOI:
10.1093/bioinformatics/btz338
复制
发表时间:
2019-07-15
期刊:
影响因子:
5.8
通讯作者:
Pfeifer, Nico
Pfeifer, Nico
中科院分区:
生物学3区
文献类型:
--
作者:
Handl, Lisa;Jalali, Adrin;Pfeifer, Nico

文献摘要

被引文献

相似文献

预测模型是解决计算生物学中复杂问题的有力工具。它们通常被设计用于预测或分类来自与训练数据相同的未知分布的数据。然而,在许多现实环境中,不受控制的生物或技术因素可能导致不同时间采集的数据集之间的分布不匹配,导致模型性能在新数据上恶化。计算生物学中一个常见的额外障碍是具有比样本多得多的特征的数据稀缺。为了解决这些问题,我们提出了一种基于加权弹性网络的无监督域自适应方法。我们方法的关键思想是比较训练和测试数据中输入之间的依赖关系,并增加弹性网络正则化项中不同行为特征的成本。在这样做时,我们鼓励该模型分配一个更高的重要性,是强大的和跨domain.Results行为相似的功能,我们评估我们的方法与不同程度的分布不匹配的模拟数据和真实的数据,考虑在多个组织的DNA甲基化数据的基础上的年龄预测的问题。与非自适应标准模型相比,我们的方法大大减少了样本与不匹配的分布的错误。在真实的数据上,我们在小脑样本上实现了低得多的误差,小脑组织不是训练数据的一部分,并且标准模型预测得很差。我们的研究结果表明,无监督域自适应是可能的计算生物学中的应用程序,即使有更多的功能比samples.Availability和实现源代码可在https://github.com/PfeiferLabTue/wenda.Supplementary信息补充数据可在生物信息学在线。
Motivation Predictive models are a powerful tool for solving complex problems in computational biology. They are typically designed to predict or classify data coming from the same unknown distribution as the training data. In many real-world settings, however, uncontrolled biological or technical factors can lead to a distribution mismatch between datasets acquired at different times, causing model performance to deteriorate on new data. A common additional obstacle in computational biology is scarce data with many more features than samples. To address these problems, we propose a method for unsupervised domain adaptation that is based on a weighted elastic net. The key idea of our approach is to compare dependencies between inputs in training and test data and to increase the cost of differently behaving features in the elastic net regularization term. In doing so, we encourage the model to assign a higher importance to features that are robust and behave similarly across domains.Results We evaluate our method both on simulated data with varying degrees of distribution mismatch and on real data, considering the problem of age prediction based on DNA methylation data across multiple tissues. Compared with a non-adaptive standard model, our approach substantially reduces errors on samples with a mismatched distribution. On real data, we achieve far lower errors on cerebellum samples, a tissue which is not part of the training data and poorly predicted by standard models. Our results demonstrate that unsupervised domain adaptation is possible for applications in computational biology, even with many more features than samples.Availability and implementation Source code is available at https://github.com/PfeiferLabTue/wenda.Supplementary informationSupplementary data are available at Bioinformatics online.