Hadoop Recognition of Biomedical Named Entity Using Conditional Random Fields

Hadoop Recognition of Biomedical Named Entity Using Conditional Random Fields
复制标题

使用条件随机字段的 Hadoop 识别生物医学命名实体

DOI:
10.1109/tpds.2014.2368568
复制
发表时间:
2015-11-01
影响因子:
5.3
通讯作者:
Hwang, Kai
Hwang, Kai
中科院分区:
计算机科学2区
文献类型:
--
作者:
Li, Kenli;Ai, Wei;Hwang, Kai

文献摘要

被引文献

相似文献

处理大量数据提出了一个具有挑战性的问题,尤其是在数据冗余系统中。作为最知名的模型之一,条件随机场(CRF)模型已被广泛应用于生物医学命名实体识别(Bio-ner)。由于内部顺序的特征,CRF模型的性能提高是非平凡的,需要新的并行解决方案。通过将有限的内存Broyden-fletcher-Goldfarb-Shanno(L-BFGS)和Viterbi算法结合在一起,我们在本文中提出了一种称为MapReduce CRF(MRCRF)的平行CRF算法,其中包含两个并行子词法,以处理两个时空的crff of the crff of crff。 MAPREDUCE L-BFGS(MRLB)算法利用MapReduce框架来增强估计参数的能力。此外,MapReduce Viterbi(MRVTB)算法通过使用另一个MapReduce作业扩展了Viterbi算法,从而扩展了最有可能的状态序列。实验结果表明,MRCRF算法通过在时间效率方面表现出显着的性能提高并保留保证的正确性水平,从而优于其他竞争方法。
Processing large volumes of data has presented a challenging issue, particularly in data-redundant systems. As one of the most recognized models, the conditional random fields (CRF) model has been widely applied in biomedical named entity recognition (Bio-NER). Due to the internally sequential feature, performance improvement of the CRF model is nontrivial, which requires new parallelized solutions. By combining and parallelizing the limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) and Viterbi algorithms, we propose a parallel CRF algorithm called MapReduce CRF (MRCRF) in this paper, which contains two parallel sub-algorithms to handle two time-consuming steps of the CRF model. The MapReduce L-BFGS (MRLB) algorithm leverages the MapReduce framework to enhance the capability of estimating parameters. Furthermore, the MapReduce Viterbi (MRVtb) algorithm infers the most likely state sequence by extending the Viterbi algorithm with another MapReduce job. Experimental results show that the MRCRF algorithm outperforms other competing methods by exhibiting significant performance improvement in terms of time efficiency as well as preserving a guaranteed level of correctness.