A global network of biomedical relationships derived from text.

A global network of biomedical relationships derived from text.
复制标题

DOI:
10.1093/bioinformatics/bty114
复制
发表时间:
2018-08-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Altman RB
Altman RB
中科院分区:
其他
文献类型:
--
作者:
Percha B;Altman RB

文献摘要

参考文献

被引文献

相似文献

生物医学界对化学物质、基因和表型如何相互作用的集体理解分布在超过 2400 万篇研究文章的文本中。这些相互作用提供了对高阶生化现象背后机制的见解,例如药物间相互作用和个体之间药物反应的变化。为了帮助他们大规模管理,我们必须了解哪些关系类型是可能的,并将非结构化自然语言描述映射到这些结构化类上。我们使用 NCBI 的 PubTator 注释来识别 Medline 摘要中的化学、基因和疾病名称实例,并应用斯坦福依赖解析器来查找单个句子中实体对之间的连接依赖路径。我们将已发布的集成双聚类算法(EBC)与层次聚类相结合,将依赖路径分组为语义相关的类别,并用标签或“主题”(例如“抑制”和“激活”)进行注释。我们根据六个人工数据库评估了我们的主题分配:DrugBank、Reactome、SIDER、治疗目标数据库、OMIM 和 PharmGKB。聚类揭示了化学-基因关系的 10 个广泛主题,7 个化学-疾病主题,10 个基因-疾病主题,9 个基因-基因关系主题。在大多数情况下,丰富的主题直接对应于已知的数据库关系。我们的最终数据集以网络形式表示,包含 37 491 个主题标记的化学基因边、2 021 192 个化学疾病边、136 206 个基因-疾病边和 41 418 个基因-基因边,每个边代表文献中某处相互作用的单句描述。完整的网络可在 Zenodo (https://zenodo.org/record/1035500) 上找到。我们还提供了连接 Medline 摘要中的生物医学实体的全套依赖路径以及相关句子,以供生物医学研究界将来使用。 补充数据可在生物信息学在线获取。
The biomedical community’s collective understanding of how chemicals, genes and phenotypes interact is distributed across the text of over 24 million research articles. These interactions offer insights into the mechanisms behind higher order biochemical phenomena, such as drug-drug interactions and variations in drug response across individuals. To assist their curation at scale, we must understand what relationship types are possible and map unstructured natural language descriptions onto these structured classes. We used NCBI’s PubTator annotations to identify instances of chemical, gene and disease names in Medline abstracts and applied the Stanford dependency parser to find connecting dependency paths between pairs of entities in single sentences. We combined a published ensemble biclustering algorithm (EBC) with hierarchical clustering to group the dependency paths into semantically-related categories, which we annotated with labels, or ‘themes’ (‘inhibition’ and ‘activation’, for example). We evaluated our theme assignments against six human-curated databases: DrugBank, Reactome, SIDER, the Therapeutic Target Database, OMIM and PharmGKB. Clustering revealed 10 broad themes for chemical-gene relationships, 7 for chemical-disease, 10 for gene-disease and 9 for gene–gene. In most cases, enriched themes corresponded directly to known database relationships. Our final dataset, represented as a network, contained 37 491 thematically-labeled chemical-gene edges, 2 021 192 chemical-disease edges, 136 206 gene-disease edges and 41 418 gene–gene edges, each representing a single-sentence description of an interaction from somewhere in the literature. The complete network is available on Zenodo (https://zenodo.org/record/1035500). We have also provided the full set of dependency paths connecting biomedical entities in Medline abstracts, with associated sentences, for future use by the biomedical research community. Supplementary data are available at Bioinformatics online.
DOI: 10.1016/j.jbi.2009.02.002
发表时间: 2009-04
影响因子: 4.5
作者:
Cohen T;Widdows D
通讯作者: Widdows D
DOI: 10.1097/00008571-200409000-00002
发表时间: 2004-09-01
期刊: PHARMACOGENETICS
影响因子: --
作者:
Chang, JT;Altman, RB
通讯作者: Altman, RB
DOI: 10.1198/jasa.2011.tm10183
发表时间: 2011-09-01
影响因子: 3.7
作者:
Bien, Jacob;Tibshirani, Robert
通讯作者: Tibshirani, Robert
DOI: 10.1016/j.jbi.2010.07.006
发表时间: 2011-02
影响因子: 4.5
作者:
Liu K;Hogan WR;Crowley RS
通讯作者: Crowley RS
DOI: 10.1093/bioinformatics/btv476
发表时间: 2016-01-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Mallory EK;Zhang C;Ré C;Altman RB
通讯作者: Altman RB