FamPlex: a resource for entity recognition and relationship resolution of human protein families and complexes in biomedical text mining.

FamPlex: a resource for entity recognition and relationship resolution of human protein families and complexes in biomedical text mining.
复制标题

DOI:
10.1186/s12859-018-2211-5
复制
发表时间:
2018-06-28
期刊:
影响因子:
3
通讯作者:
Sorger PK
Sorger PK
中科院分区:
生物学4区
文献类型:
--
作者:
Bachman JA;Gyori BM;Sorger PK

文献摘要

参考文献

被引文献

相似文献

为了自动阅读科学出版物以提取关于分子机制的有用信息,关键是基因、蛋白质和其他实体与统一标识符正确关联,该过程被称为命名实体链接或“接地”。正确的基础对于解决挖掘的信息,策划的交互数据库和生物数据集之间的关系至关重要。这一过程的准确性在很大程度上取决于机器可读资源的可用性,这些资源将生物医学文献中常见的同义词和缩写与统一的标识符相关联。在一项涉及使用REACH事件提取软件自动阅读215,000篇文章的任务中,我们发现多蛋白质家族的基础不成比例地不准确(例如,“AKT”)和具有多个亚基的复合物(例如,“NF- κB”)。为了解决这个问题,我们构建了一个人工策划的资源,定义蛋白质家族和复合物,因为它们通常在生物医学文本中遇到。在Families中,家族和复合体的基因水平成分以灵活的格式定义,允许多层次的成员资格。为了创建Families,根据经验从文献中识别与实体对应的文本字符串,并手动链接到统一标识符;这些标识符也映射到多个相关数据库中的等效条目。Families还包括精选的前缀和后缀模式,可改进命名实体识别和事件提取。在一个包含54,000篇文章的测试语料库中对REACH提取的评估表明,Familytree显着提高了家族和复合体的基础准确性(从15%到71%)。FAMILY中实体的层次结构也使得整合家族、亚家族和单个蛋白质之间原本不相关的机制信息成为可能。将Famalgins应用于TRIPS/DRUM阅读系统和Biocreative VI生物实体标准化任务数据集,证明了Famalgins在其他环境中的实用性。FAMILY是一个有效的资源,用于提高命名实体识别,接地,并在生物医学文本的自动阅读的关系解析。FAMILY中的内容在Creative Commons CC 0许可下以表格和开放生物医学本体的格式在https://github.com/sorgerlab/famplex上提供,并已集成到TRIPS/DRUM和REACH阅读系统中。
For automated reading of scientific publications to extract useful information about molecular mechanisms it is critical that genes, proteins and other entities be correctly associated with uniform identifiers, a process known as named entity linking or “grounding.” Correct grounding is essential for resolving relationships among mined information, curated interaction databases, and biological datasets. The accuracy of this process is largely dependent on the availability of machine-readable resources associating synonyms and abbreviations commonly found in biomedical literature with uniform identifiers. In a task involving automated reading of ∼215,000 articles using the REACH event extraction software we found that grounding was disproportionately inaccurate for multi-protein families (e.g., “AKT”) and complexes with multiple subunits (e.g.“NF- κB”). To address this problem we constructed FamPlex, a manually curated resource defining protein families and complexes as they are commonly encountered in biomedical text. In FamPlex the gene-level constituents of families and complexes are defined in a flexible format allowing for multi-level, hierarchical membership. To create FamPlex, text strings corresponding to entities were identified empirically from literature and linked manually to uniform identifiers; these identifiers were also mapped to equivalent entries in multiple related databases. FamPlex also includes curated prefix and suffix patterns that improve named entity recognition and event extraction. Evaluation of REACH extractions on a test corpus of ∼54,000 articles showed that FamPlex significantly increased grounding accuracy for families and complexes (from 15 to 71%). The hierarchical organization of entities in FamPlex also made it possible to integrate otherwise unconnected mechanistic information across families, subfamilies, and individual proteins. Applications of FamPlex to the TRIPS/DRUM reading system and the Biocreative VI Bioentity Normalization Task dataset demonstrated the utility of FamPlex in other settings. FamPlex is an effective resource for improving named entity recognition, grounding, and relationship resolution in automated reading of biomedical text. The content in FamPlex is available in both tabular and Open Biomedical Ontology formats at https://github.com/sorgerlab/famplex under the Creative Commons CC0 license and has been integrated into the TRIPS/DRUM and REACH reading systems.
DOI: 10.1038/nbt1346
发表时间: 2007-11-01
影响因子: 46.9
作者:
Smith, Barry;Ashburner, Michael;Lewis, Suzanna
通讯作者: Lewis, Suzanna
DOI: 10.1093/nar/gki031
发表时间: 2005-01-01
影响因子: 14.9
作者:
Maglott D;Ostell J;Pruitt KD;Tatusova T
通讯作者: Tatusova T
DOI: 10.1093/bfgp/elu015
发表时间: 2015-05
影响因子: 4
作者:
Ananiadou S;Thompson P;Nawaz R;McNaught J;Kell DB
通讯作者: Kell DB
DOI: 10.7554/elife.04640
发表时间: 2015-08-18
期刊: eLife
影响因子: 7.7
作者:
Korkut A;Wang W;Demir E;Aksoy BA;Jing X;Molinelli EJ;Babur Ö;Bemis DL;Onur Sumer S;Solit DB;Pratilas CA;Sander C
通讯作者: Sander C
DOI: 10.1093/nar/gkq1039
发表时间: 2011-01
影响因子: 14.9
作者:
Cerami EG;Gross BE;Demir E;Rodchenkov I;Babur O;Anwar N;Schultz N;Bader GD;Sander C
通讯作者: Sander C