Embeddings of Label Components for Sequence Labeling: A Case Study of Fine-grained Named Entity Recognition

Embeddings of Label Components for Sequence Labeling: A Case Study of Fine-grained Named Entity Recognition
复制标题

DOI:
10.18653/v1/2020.acl-srw.30
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Takuma Kato;Kaori Abe;Hiroki Ouchi;Shumpei Miyawaki;Jun Suzuki;Kentaro Inui
Takuma Kato;Kaori Abe;Hiroki Ouchi;Shumpei Miyawaki;Jun Suzuki;Kentaro Inui
中科院分区:
其他
文献类型:
--
作者:
Takuma Kato;Kaori Abe;Hiroki Ouchi;Shumpei Miyawaki;Jun Suzuki;Kentaro Inui

文献摘要

相似文献

通常,序列标记中使用的标记由不同类型的元素组成。例如,IOB格式的实体标签,如B-Person和I-Person,可以分解为SPAN(B和I)和类型信息(Person)。然而,尽管大多数序列标签模型不考虑这样的标签分量,但是标签之间的共享分量,例如Person,对于标签预测是有益的。在这项工作中,我们建议将标签组件信息作为嵌入到模型中。通过在英语和日语细粒度命名实体识别上的实验,我们证明了该方法提高了性能,特别是对于具有低频标签的实例。
In general, the labels used in sequence labeling consist of different types of elements. For example, IOB-format entity labels, such as B-Person and I-Person, can be decomposed into span (B and I) and type information (Person). However, while most sequence labeling models do not consider such label components, the shared components across labels, such as Person, can be beneficial for label prediction. In this work, we propose to integrate label component information as embeddings into models. Through experiments on English and Japanese fine-grained named entity recognition, we demonstrate that the proposed method improves performance, especially for instances with low-frequency labels.