Gene Prioritization by Compressive Data Fusion and Chaining.

Gene Prioritization by Compressive Data Fusion and Chaining.
复制标题

通过压缩数据融合和链接的基因优先次序。

DOI:
10.1371/journal.pcbi.1004552
复制
发表时间:
2015-10
影响因子:
4.3
通讯作者:
Zupan B
Zupan B
中科院分区:
生物学2区
文献类型:
--
作者:
Žitnik M;Nam EA;Dinh C;Kuspa A;Shaulsky G;Zupan B

文献摘要

被引文献

相似文献

数据集成程序将异构数据集组合成预测模型,但仅限于与目标对象类型(例如基因)明确相关的数据。拼贴是一种新的基因优先级数据融合方法。它考虑与预测任务的各种关联级别的数据集,利用集体矩阵分解来压缩数据,并链接以关联数据概要中包含的不同对象类型。拼贴画根据基因与几个种子基因的相似性对基因进行优先排序。我们通过优先考虑盘基网柄菌中的细菌反应基因作为原核生物-真核生物相互作用的新模型系统来测试 Collage。 Collage 使用 4 个种子基因和 14 个数据集(其中只有一个与细菌反应直接相关)提出了 8 个候选基因,这些基因很容易被验证为盘基网柄菌对革兰氏阴性细菌反应所必需的。这些发现将拼贴确立为一种从异构且粗略相关的数据集的整合中推断生物知识的方法。在日常生活中,我们通过考虑所有可用信息来做出决定,并且经常发现即使是看似间接的证据也能提供优势。我们的新计算方法 Collage 对来自大量异构数据的基因进行优先排序。在社会性变形虫盘基网柄菌的案例研究中,我们从四个细菌反应基因和 14 个不同的数据集(从基因表达到通路和文献信息)开始。 Collage 提出了八个候选基因,并在湿实验室中进行了测试。所有八个候选者的突变都降低了变形虫在革兰氏阴性细菌上生长的能力。此外,八个候选基因中的五个是革兰氏阴性细菌生长所必需的,但对革兰氏阳性细菌的生长没有明显影响。这是一个非常准确的结果,因为据估计 12,000 个盘基网柄菌基因中只有大约 100 个基因负责细菌反应。
Data integration procedures combine heterogeneous data sets into predictive models, but they are limited to data explicitly related to the target object type, such as genes. Collage is a new data fusion approach to gene prioritization. It considers data sets of various association levels with the prediction task, utilizes collective matrix factorization to compress the data, and chaining to relate different object types contained in a data compendium. Collage prioritizes genes based on their similarity to several seed genes. We tested Collage by prioritizing bacterial response genes in Dictyostelium as a novel model system for prokaryote-eukaryote interactions. Using 4 seed genes and 14 data sets, only one of which was directly related to the bacterial response, Collage proposed 8 candidate genes that were readily validated as necessary for the response of Dictyostelium to Gram-negative bacteria. These findings establish Collage as a method for inferring biological knowledge from the integration of heterogeneous and coarsely related data sets. In everyday life, we make decisions by considering all the available information, and often find that inclusion of even seemingly circumstantial evidence provides an advantage. Our new computational method Collage prioritizes genes from a large collection of heterogeneous data. In a case study on social amoeba Dictyostelium, we started from four bacterial response genes and 14 different data sets ranging from gene expression to pathway and literature information. Collage proposed eight candidate genes that were tested in the wet laboratory. Mutations in all eight candidates reduced the ability of the amoebae to grow on Gram-negative bacteria. Furthermore, five out of the eight candidate genes were required for growth on Gram-negative bacteria but had no discernible effect on growth on Gram-positive bacteria. This is a remarkably accurate result since only about a hundred of the 12,000 Dictyostelium genes are estimated to be responsible for bacterial response.