PolySearch: a web-based text mining system for extracting relationships between human diseases, genes, mutations, drugs and metabolites.

PolySearch: a web-based text mining system for extracting relationships between human diseases, genes, mutations, drugs and metabolites.
复制标题

DOI:
10.1093/nar/gkn296
复制
发表时间:
2008-07-01
影响因子:
14.9
通讯作者:
Wishart DS
Wishart DS
中科院分区:
生物学2区
文献类型:
--
作者:
Cheng D;Knox C;Young N;Stothard P;Damaraju S;Wishart DS

文献摘要

参考文献

被引文献

相似文献

生物医学文本挖掘中的一个特殊挑战是找到处理“综合”或“关联”查询的方法,例如“查找与乳腺癌相关的所有基因”。考虑到基因组学、蛋白质组学或代谢组学中的许多查询涉及这些类型的综合搜索,我们认为可以支持这些搜索的基于网络的工具将是非常有用的。为了满足这一需求,我们开发了PolySearch Web服务器。PolySearch支持对近12种不同类型的文本、科学摘要或生物信息数据库进行超过50种不同类别的查询。PolySearch支持的典型查询是“给定X,查找所有Y”,其中X或Y可以是疾病、组织、细胞区室、基因/蛋白质名称、SNP、突变、药物和代谢物。PolySearch还利用文本挖掘和信息检索中的各种技术来识别,突出显示和排名信息摘要,段落或句子。PolySearch的性能已经在基因同义词识别、蛋白质-蛋白质相互作用识别和疾病基因识别等任务中使用各种手动组装的“金标准”文本语料库进行了评估。它在这些任务上的f-测量值分别为88%、81%和79%。这些值比其他已发布的工具好5%到50%。该服务器可在http://wishart.biology.ualberta.ca/polysearch上免费获得
A particular challenge in biomedical text mining is to find ways of handling ‘comprehensive’ or ‘associative’ queries such as ‘Find all genes associated with breast cancer’. Given that many queries in genomics, proteomics or metabolomics involve these kind of comprehensive searches we believe that a web-based tool that could support these searches would be quite useful. In response to this need, we have developed the PolySearch web server. PolySearch supports >50 different classes of queries against nearly a dozen different types of text, scientific abstract or bioinformatic databases. The typical query supported by PolySearch is ‘Given X, find all Y's’ where X or Y can be diseases, tissues, cell compartments, gene/protein names, SNPs, mutations, drugs and metabolites. PolySearch also exploits a variety of techniques in text mining and information retrieval to identify, highlight and rank informative abstracts, paragraphs or sentences. PolySearch's performance has been assessed in tasks such as gene synonym identification, protein–protein interaction identification and disease gene identification using a variety of manually assembled ‘gold standard’ text corpuses. Its f-measure on these tasks is 88, 81 and 79%, respectively. These values are between 5 and 50% better than other published tools. The server is freely available at http://wishart.biology.ualberta.ca/polysearch
HMDB:人类代谢组数据库。
DOI: 10.1093/nar/gkl923
发表时间: 2007-01
影响因子: 14.9
作者:
Wishart, David S;Tzur, Dan;Knox, Craig;Eisner, Roman;Guo, An Chi;Young, Nelson;Cheng, Dean;Jewell, Kevin;Arndt, David;Sawhney, Summit;Fung, Chris;Nikolai, Lisa;Lewis, Mike;Coutouly, Marie-Aude;Forsythe, Ian;Tang, Peter;Shrivastava, Savita;Jeroncic, Kevin;Stothard, Paul;Amegbey, Godwin;Block, David;Hau, David D;Wagner, James;Miniaci, Jessica;Clements, Melisa;Gebremedhin, Mulu;Guo, Natalie;Zhang, Ying;Duggan, Gavin E;Macinnis, Glen D;Weljie, Alim M;Dowlatabadi, Reza;Bamforth, Fiona;Clive, Derrick;Greiner, Russ;Li, Liang;Marrie, Tom;Sykes, Brian D;Vogel, Hans J;Querengesser, Lori
通讯作者: Querengesser, Lori
DOI: 10.1021/pr0340227
发表时间: 2003-07-01
影响因子: 4.4
作者:
Hu, YH;Hines, LM;LaBaer, J
通讯作者: LaBaer, J
DOI: 10.1093/bioinformatics/btl408
发表时间: 2006-10-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Plake, Conrad;Schiemann, Torsten;Leser, Ulf
通讯作者: Leser, Ulf
DOI: 10.1038/sj.onc.1203335
发表时间: 1999-12-23
期刊: ONCOGENE
影响因子: 8
作者:
Baasiri, RA;Glasser, SR;Wheeler, DA
通讯作者: Wheeler, DA
DOI: 10.1093/bioinformatics/bti493
发表时间: 2005-08-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hao, Y;Zhu, XY;Li, M
通讯作者: Li, M