Deep learning with word embeddings improves biomedical named entity recognition.

Deep learning with word embeddings improves biomedical named entity recognition.
复制标题

DOI:
10.1093/bioinformatics/btx228
复制
发表时间:
2017-07-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Leser U
Leser U
中科院分区:
其他
文献类型:
--
作者:
Habibi M;Weber L;Neves M;Wiegandt DL;Leser U

文献摘要

参考文献

被引文献

相似文献

文本挖掘已经成为生物医学研究的重要工具。最基本的文本挖掘任务是识别生物医学命名实体(NER),如基因,化学物质和疾病。目前的NER方法依赖于预定义的功能,试图捕捉特定的表面属性的实体类型,典型的本地上下文,背景知识和语言信息的属性。最先进的工具是特定于实体的,因为字典和经验最佳特征集在实体类型之间不同,这使得它们的开发成本很高。此外,特征通常针对特定的黄金标准语料库进行优化,这使得质量度量的外推变得困难。我们证明了一种基于深度学习和统计词嵌入的完全通用的方法[称为长短期记忆网络条件随机场(LSTM-CRF)]优于最先进的实体特定NER工具,并且通常是大幅度的。为此,我们比较了LSTM-CRF在33个数据集上的性能,这些数据集覆盖了5个不同的实体类,并与一流的NER工具和实体不可知的CRF实现进行了比较。平均而言,LSTM-CRF的F1分数比基线高5%,主要是由于回忆的急剧增加。LSTM-CRF的源代码可在https://github.com/glample/tagger上获得,语料库的链接可在https://corposaurus.github.io/corpora/上获得。
Text mining has become an important tool for biomedical research. The most fundamental text-mining task is the recognition of biomedical named entities (NER), such as genes, chemicals and diseases. Current NER methods rely on pre-defined features which try to capture the specific surface properties of entity types, properties of the typical local context, background knowledge, and linguistic information. State-of-the-art tools are entity-specific, as dictionaries and empirically optimal feature sets differ between entity types, which makes their development costly. Furthermore, features are often optimized for a specific gold standard corpus, which makes extrapolation of quality measures difficult. We show that a completely generic method based on deep learning and statistical word embeddings [called long short-term memory network-conditional random field (LSTM-CRF)] outperforms state-of-the-art entity-specific NER tools, and often by a large margin. To this end, we compared the performance of LSTM-CRF on 33 data sets covering five different entity classes with that of best-of-class NER tools and an entity-agnostic CRF implementation. On average, F1-score of LSTM-CRF is 5% above that of the baselines, mostly due to a sharp increase in recall. The source code for LSTM-CRF is available at https://github.com/glample/tagger and the links to the corpora are available at https://corposaurus.github.io/corpora/.
DOI: 10.1186/s13321-016-0172-0
发表时间: 2016
影响因子: 8.6
作者:
Habibi M;Wiegandt DL;Schmedding F;Leser U
通讯作者: Leser U
DOI: 10.1186/1758-2946-7-s1-s6
发表时间: 2015
影响因子: 8.6
作者:
Batista-Navarro R;Rak R;Ananiadou S
通讯作者: Ananiadou S
DOI: 10.1016/j.jbi.2013.12.006
发表时间: 2014-02
影响因子: 4.5
作者:
Dogan, Rezarta Islamaj;Leaman, Robert;Lu, Zhiyong
通讯作者: Lu, Zhiyong
DOI: 10.1371/journal.pone.0107477
发表时间: 2014
期刊: PloS one
影响因子: 3.7
作者:
Akhondi SA;Klenner AG;Tyrchan C;Manchala AK;Boppana K;Lowe D;Zimmermann M;Jagarlapudi SA;Sayle R;Kors JA;Muresan S
通讯作者: Muresan S
DOI: 10.1016/j.neunet.2005.06.042
发表时间: 2005-06-01
期刊: NEURAL NETWORKS
影响因子: 7.8
作者:
Graves, A;Schmidhuber, J
通讯作者: Schmidhuber, J