ExCAPE-DB: an integrated large scale dataset facilitating Big Data analysis in chemogenomics.

ExCAPE-DB: an integrated large scale dataset facilitating Big Data analysis in chemogenomics.
复制标题

DOI:
10.1186/s13321-017-0203-5
复制
发表时间:
2017
影响因子:
8.6
通讯作者:
Chen H
Chen H
中科院分区:
化学2区
文献类型:
--
作者:
Sun J;Jeliazkova N;Chupakin V;Golib-Dzib JF;Engkvist O;Carlsson L;Wegner J;Ceulemans H;Georgiev I;Jeliazkov V;Kochev N;Ashby TJ;Chen H

文献摘要

被引文献

相似文献

化学基因组学数据通常是指化合物对蛋白质靶点阵列的活性数据,并且代表用于构建计算机靶点预测模型的重要信息来源。越来越多的化学基因组学数据为基于大数据构建模型提供了令人兴奋的机会。准备一个高质量的数据集是实现这一目标的重要一步,这项工作的目的是编译这样一个全面的化学基因组学数据集。该数据集包括来自公共数据库(PubChem和ChEMBL)的7000多万个SAR数据点,包括结构,目标信息和活动注释。我们的愿望是创建一个有用的化学基因组学资源,反映工业规模的数据,不仅用于建立计算机多药理学和脱靶效应的预测模型,而且用于验证一般的化学信息学方法。本文的在线版本(doi:10.1186/s13321-017-0203-5)包含补充材料,可供授权用户使用。
Chemogenomics data generally refers to the activity data of chemical compounds on an array of protein targets and represents an important source of information for building in silico target prediction models. The increasing volume of chemogenomics data offers exciting opportunities to build models based on Big Data. Preparing a high quality data set is a vital step in realizing this goal and this work aims to compile such a comprehensive chemogenomics dataset. This dataset comprises over 70 million SAR data points from publicly available databases (PubChem and ChEMBL) including structure, target information and activity annotations. Our aspiration is to create a useful chemogenomics resource reflecting industry-scale data not only for building predictive models of in silico polypharmacology and off-target effects but also for the validation of cheminformatics approaches in general. The online version of this article (doi:10.1186/s13321-017-0203-5) contains supplementary material, which is available to authorized users.