NERO: a biomedical named-entity (recognition) ontology with a large, annotated corpus reveals meaningful associations through text embedding.

NERO: a biomedical named-entity (recognition) ontology with a large, annotated corpus reveals meaningful associations through text embedding.
复制标题

DOI:
10.1038/s41540-021-00200-x
复制
发表时间:
2021-10-20
影响因子:
4
通讯作者:
Rzhetsky A
Rzhetsky A
中科院分区:
生物学2区
文献类型:
--
作者:
Wang K;Stevens R;Alachram H;Li Y;Soldatova L;King R;Ananiadou S;Schoene AM;Li M;Christopoulou F;Ambite JL;Matthew J;Garg S;Hermjakob U;Marcu D;Sheng E;Beißbarth T;Wingender E;Galstyan A;Gao X;Chambers B;Pan W;Khomtchouk BB;Evans JA;Rzhetsky A

文献摘要

参考文献

相似文献

机器阅读(MR)对于解锁数百万现有生物医学文档中包含的有价值的知识至关重要。在过去的二十年里,最引人注目的进展,在MR紧随其后的关键语料库的发展。大型、注释良好的语料库与MR方法和自动知识提取系统的不断进步有关,就像ImageNet是开发机器视觉技术的基础一样。本研究为一个先进的生物医学命名实体分析工具贡献了六个组成部分:(a)一个新的命名实体识别本体(NERO),专门用于描述生物医学文本中的文本实体,它解释了不同层次的歧义,桥接了分子生物学,遗传学,生物化学和医学的科学子语言;(B)人类专家注释数百个命名实体类的详细指南;(c)所有命名实体的象形图,以简化管理员的注释负担;(d)一个原始的、带注释的语料库,包括35,865个句子,其中封装了190,679个命名实体和43,438个连接两个或多个实体的事件;(f)嵌入模型,展示嵌入该语料库中的生物医学协会的前景。
Machine reading (MR) is essential for unlocking valuable knowledge contained in millions of existing biomedical documents. Over the last two decades, the most dramatic advances in MR have followed in the wake of critical corpus development. Large, well-annotated corpora have been associated with punctuated advances in MR methodology and automated knowledge extraction systems in the same way that ImageNet was fundamental for developing machine vision techniques. This study contributes six components to an advanced, named entity analysis tool for biomedicine: (a) a new, Named Entity Recognition Ontology (NERO) developed specifically for describing textual entities in biomedical texts, which accounts for diverse levels of ambiguity, bridging the scientific sublanguages of molecular biology, genetics, biochemistry, and medicine; (b) detailed guidelines for human experts annotating hundreds of named entity classes; (c) pictographs for all named entities, to simplify the burden of annotation for curators; (d) an original, annotated corpus comprising 35,865 sentences, which encapsulate 190,679 named entities and 43,438 events connecting two or more entities; (e) validated, off-the-shelf, named entity recognition (NER) automated extraction, and; (f) embedding models that demonstrate the promise of biomedical associations embedded within this corpus.
DOI: 10.1073/pnas.1720347115
发表时间: 2018-04-17
影响因子: 11.1
作者:
Garg, Nikhil;Schiebinger, Londa;Zou, James
通讯作者: Zou, James
DOI: 10.7717/peerj-cs.644
发表时间: 2021
期刊: PeerJ. Computer science
影响因子: --
作者:
Kwak H;An J;Jing E;Ahn YY
通讯作者: Ahn YY
DOI: 10.1016/j.jbi.2013.12.006
发表时间: 2014-02
影响因子: 4.5
作者:
Dogan, Rezarta Islamaj;Leaman, Robert;Lu, Zhiyong
通讯作者: Lu, Zhiyong
DOI: 10.1007/s100510050276
发表时间: 1998-04-01
影响因子: 1.6
作者:
Laherrere, J;Sornette, D
通讯作者: Sornette, D
DOI: 10.1126/science.aal4230
发表时间: 2017-04-14
期刊: SCIENCE
影响因子: 56.9
作者:
Caliskan, Aylin;Bryson, Joanna J.;Narayanan, Arvind
通讯作者: Narayanan, Arvind