Triage by ranking to support the curation of protein interactions.

Triage by ranking to support the curation of protein interactions.
复制标题

通过排名来支撑蛋白质相互作用的策划来分类。

DOI:
10.1093/database/bax040
复制
发表时间:
2017-01-01
期刊:
Database : the journal of biological databases and curation
影响因子:
--
通讯作者:
Ruch P
Ruch P
中科院分区:
其他
文献类型:
--
作者:
Mottin L;Pasche E;Gobeill J;Rech de Laval V;Gleizes A;Michel PA;Bairoch A;Gaudet P;Ruch P

文献摘要

参考文献

被引文献

相似文献

今天,分子生物学数据库是生命和健康科学知识共享的基石。这些资源的管理和维护是劳动密集型的。尽管文本挖掘在策展人中获得了推动力,但它在策展工作流中的集成还没有被广泛采用。瑞士生物信息学研究所的文本挖掘和CALIPHO小组联手设计了一个新的管理支持系统,名为NextA5。在这份报告中,我们探索了新的分类服务的集成,以支持两种类型的生物数据的管理:蛋白质-蛋白质相互作用(PPI)和翻译后修饰(PTM)。PPI和PTMS的识别是一项特殊的挑战,因为它不仅需要识别生物实体(蛋白质或残基),而且还需要识别特定的关系(例如结合或位置)。这些关系不能用诸如分子功能的基因本体论之类的本体描述符来描述,这使得分类任务更具挑战性。因此,为这些任务确定论文的优先次序需要制定不同的方法。在这份报告中,我们提出了一种新的方法来对包含特定于PPI和PTM的信息的文章进行优先排序。新的资源(RESTful API、带语义注释的MEDLINE库)丰富了neXtA5平台。我们在CALIPHO小组之前注释的一组100个蛋白质上调整了文章优先顺序模型。使用200个带注释的蛋白质的数据集测试了分类服务的有效性。我们定义了两组描述符来支持自动分拣:第一组用于丰富具有PPI数据的论文,第二组用于PTMS。这些描述符的所有出现都在MEDLINE中进行了标记和索引,从而构成了MEDLINE的语义注释版本。然后使用这些注释来估计特定文章与所选注释类型的相关性。该相关性分数与本地向量空间搜索引擎相结合,以生成PMID的排名列表。我们还评估了一种查询优化策略,该策略将特定的关键字(如‘绑定’或‘交互’)添加到原始查询中。与PubMed相比,nextA5分流服务对具有PPI信息的论文进行优先排序的搜索效率提高了190%,对于具有PTMS信息的论文的搜索效率提高了260%。将高级检索和查询优化策略与自动丰富的MEDLINE内容相结合,可以有效地改进复杂检索任务中的分类,例如蛋白质PPI和PTM的检索。数据库地址:http://candy.hesge.ch/nextA5
Today, molecular biology databases are the cornerstone of knowledge sharing for life and health sciences. The curation and maintenance of these resources are labour intensive. Although text mining is gaining impetus among curators, its integration in curation workflow has not yet been widely adopted. The Swiss Institute of Bioinformatics Text Mining and CALIPHO groups joined forces to design a new curation support system named nextA5. In this report, we explore the integration of novel triage services to support the curation of two types of biological data: protein–protein interactions (PPIs) and post-translational modifications (PTMs). The recognition of PPIs and PTMs poses a special challenge, as it not only requires the identification of biological entities (proteins or residues), but also that of particular relationships (e.g. binding or position). These relationships cannot be described with onto-terminological descriptors such as the Gene Ontology for molecular functions, which makes the triage task more challenging. Prioritizing papers for these tasks thus requires the development of different approaches. In this report, we propose a new method to prioritize articles containing information specific to PPIs and PTMs. The new resources (RESTful APIs, semantically annotated MEDLINE library) enrich the neXtA5 platform. We tuned the article prioritization model on a set of 100 proteins previously annotated by the CALIPHO group. The effectiveness of the triage service was tested with a dataset of 200 annotated proteins. We defined two sets of descriptors to support automatic triage: the first set to enrich for papers with PPI data, and the second for PTMs. All occurrences of these descriptors were marked-up in MEDLINE and indexed, thus constituting a semantically annotated version of MEDLINE. These annotations were then used to estimate the relevance of a particular article with respect to the chosen annotation type. This relevance score was combined with a local vector-space search engine to generate a ranked list of PMIDs. We also evaluated a query refinement strategy, which adds specific keywords (such as ‘binds’ or ‘interacts’) to the original query. Compared to PubMed, the search effectiveness of the nextA5 triage service is improved by 190% for the prioritization of papers with PPIs information and by 260% for papers with PTMs information. Combining advanced retrieval and query refinement strategies with automatically enriched MEDLINE contents is effective to improve triage in complex curation tasks such as the curation of protein PPIs and PTMs. Database URL: http://candy.hesge.ch/nextA5
DOI: 10.1016/j.ipm.2014.10.007
发表时间: 2015-03-01
影响因子: 8.6
作者:
Chifu, Adrian-Gabriel;Hristea, Florentina;Popescu, Marius
通讯作者: Popescu, Marius
DOI: 10.1186/s12859-016-1092-8
发表时间: 2016-07-25
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Abdulla, Ahmed AbdoAziz Ahmed;Lin, Hongfei;Banbhrani, Santosh Kumar
通讯作者: Banbhrani, Santosh Kumar
DOI: 10.1145/1416950.1416952
发表时间: 2009-01-01
影响因子: 5.6
作者:
Moffat, Alistair;Zobel, Justin
通讯作者: Zobel, Justin
DOI: 10.1186/1471-2105-4-11
发表时间: 2003-03-27
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Donaldson, I;Martin, J;de Bruijn, B;Wolting, C;Lay, V;Tuekam, B;Zhang, SD;Baskin, B;Bader, GD;Michalickova, K;Pawson, T;Hogue, CWV
通讯作者: Hogue, CWV
DOI: 10.1016/j.jbi.2015.07.010
发表时间: 2015-10
影响因子: 4.5
作者:
Leaman R;Khare R;Lu Z
通讯作者: Lu Z