Metric learning on expression data for gene function prediction.

Metric learning on expression data for gene function prediction.
复制标题

DOI:
10.1093/bioinformatics/btz731
复制
发表时间:
2020-02-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
van Ham RCHJ
van Ham RCHJ
中科院分区:
其他
文献类型:
--
作者:
Makrodimitris S;Reinders MJT;van Ham RCHJ

文献摘要

参考文献

被引文献

相似文献

两个基因在不同条件下的共同表达表明它们参与了相同的生物过程。然而,当使用具有来自不同来源的许多实验条件的RNA-Seq数据集时,预计只有实验条件的子集与寻找与特定基因本体论(GO)术语相关的基因相关。因此,我们假设,当目的是寻找功能相似的基因时,不应该在所有样本上确定基因的共同表达,而应该只在那些对GO感兴趣的项有信息的样本上确定。为了解决这个问题,我们开发了协同表达的度量学习(MLC),这是一种快速算法,为每个表情样本分配特定于GO项的权重。我们的目标是获得一个加权的共表达度量,它比未加权的皮尔逊相关性更适合于应用基于关联的内疚函数预测。更具体地说,如果用给定的GO项注释两个基因,MLC会尝试最大化它们的加权共表达,此外,如果其中一个基因没有用该项注释,加权共表达就会最小化。我们在公开可用的拟南芥RNA-Seq数据上的实验表明,MLC在以术语为中心的性能上优于标准的Pearson相关。此外,我们的方法特别擅长于更具体的术语,也就是最有趣的术语。最后,通过观察特定围棋术语的样本权重,人们可以确定哪些实验对于学习该术语是重要的,并有可能确定相关的新条件,如A.thaliana和铜绿假单胞菌的实验所证明的那样。MLC以Python包的形式可在www.githorb.com/staakro/mlc上获得。补充数据可在生物信息学在线上获得。
Co-expression of two genes across different conditions is indicative of their involvement in the same biological process. However, when using RNA-Seq datasets with many experimental conditions from diverse sources, only a subset of the experimental conditions is expected to be relevant for finding genes related to a particular Gene Ontology (GO) term. Therefore, we hypothesize that when the purpose is to find similarly functioning genes, the co-expression of genes should not be determined on all samples but only on those samples informative for the GO term of interest. To address this, we developed Metric Learning for Co-expression (MLC), a fast algorithm that assigns a GO-term-specific weight to each expression sample. The goal is to obtain a weighted co-expression measure that is more suitable than the unweighted Pearson correlation for applying Guilt-By-Association-based function predictions. More specifically, if two genes are annotated with a given GO term, MLC tries to maximize their weighted co-expression and, in addition, if one of them is not annotated with that term, the weighted co-expression is minimized. Our experiments on publicly available Arabidopsis thaliana RNA-Seq data demonstrate that MLC outperforms standard Pearson correlation in term-centric performance. Moreover, our method is particularly good at more specific terms, which are the most interesting. Finally, by observing the sample weights for a particular GO term, one can identify which experiments are important for learning that term and potentially identify novel conditions that are relevant, as demonstrated by experiments in both A. thaliana and Pseudomonas Aeruginosa. MLC is available as a Python package at www.github.com/stamakro/MLC. Supplementary data are available at Bioinformatics online.
DOI: 10.1093/pcp/pcx191
发表时间: 2018-01-01
影响因子: 4.9
作者:
Obayashi T;Aoki Y;Tadaka S;Kagaya Y;Kinoshita K
通讯作者: Kinoshita K
DOI: 10.1038/s41467-018-06772-3
发表时间: 2018-10-31
影响因子: 16.6
作者:
Chen D;Yan W;Fu LY;Kaufmann K
通讯作者: Kaufmann K
DOI: 10.1093/bioinformatics/btt228
发表时间: 2013-07-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Clark WT;Radivojac P
通讯作者: Radivojac P
DOI: 10.1186/s13059-016-1037-6
发表时间: 2016-09-07
期刊: Genome biology
影响因子: 12.3
作者:
Jiang Y;Oron TR;Clark WT;Bankapur AR;D'Andrea D;Lepore R;Funk CS;Kahanda I;Verspoor KM;Ben-Hur A;Koo da CE;Penfold-Brown D;Shasha D;Youngs N;Bonneau R;Lin A;Sahraeian SM;Martelli PL;Profiti G;Casadio R;Cao R;Zhong Z;Cheng J;Altenhoff A;Skunca N;Dessimoz C;Dogan T;Hakala K;Kaewphan S;Mehryary F;Salakoski T;Ginter F;Fang H;Smithers B;Oates M;Gough J;Törönen P;Koskinen P;Holm L;Chen CT;Hsu WL;Bryson K;Cozzetto D;Minneci F;Jones DT;Chapman S;Bkc D;Khan IK;Kihara D;Ofer D;Rappoport N;Stern A;Cibrian-Uhalte E;Denny P;Foulger RE;Hieta R;Legge D;Lovering RC;Magrane M;Melidoni AN;Mutowo-Meullenet P;Pichler K;Shypitsyna A;Li B;Zakeri P;ElShal S;Tranchevent LC;Das S;Dawson NL;Lee D;Lees JG;Sillitoe I;Bhat P;Nepusz T;Romero AE;Sasidharan R;Yang H;Paccanaro A;Gillis J;Sedeño-Cortés AE;Pavlidis P;Feng S;Cejuela JM;Goldberg T;Hamp T;Richter L;Salamov A;Gabaldon T;Marcet-Houben M;Supek F;Gong Q;Ning W;Zhou Y;Tian W;Falda M;Fontana P;Lavezzo E;Toppo S;Ferrari C;Giollo M;Piovesan D;Tosatto SC;Del Pozo A;Fernández JM;Maietta P;Valencia A;Tress ML;Benso A;Di Carlo S;Politano G;Savino A;Rehman HU;Re M;Mesiti M;Valentini G;Bargsten JW;van Dijk AD;Gemovic B;Glisic S;Perovic V;Veljkovic V;Veljkovic N;Almeida-E-Silva DC;Vencio RZ;Sharan M;Vogel J;Kansakar L;Zhang S;Vucetic S;Wang Z;Sternberg MJ;Wass MN;Huntley RP;Martin MJ;O'Donovan C;Robinson PN;Moreau Y;Tramontano A;Babbitt PC;Brenner SE;Linial M;Orengo CA;Rost B;Greene CS;Mooney SD;Friedberg I;Radivojac P
通讯作者: Radivojac P
DOI: 10.1093/bioinformatics/btx143
发表时间: 2017-07-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Petryszak R;Fonseca NA;Füllgrabe A;Huerta L;Keays M;Tang YA;Brazma A
通讯作者: Brazma A