Leveraging HTML in Free Text Web Named Entity Recognition

Leveraging HTML in Free Text Web Named Entity Recognition
复制标题

DOI:
10.18653/v1/2020.coling-main.36
复制
发表时间:
2020-12
期刊:
--
影响因子:
--
通讯作者:
Colin Ashby;David Weir
Colin Ashby;David Weir
中科院分区:
其他
文献类型:
--
作者:
Colin Ashby;David Weir

文献摘要

相似文献

HTML标记通常在网页的自由文本命名实体识别中被丢弃。我们调查这些丢弃的标签是否可以用来提高NER性能。我们比较文本+标签的句子与它们的纯文本等价物,超过五个数据集,两个自由文本分割粒度和两个NER模型。我们发现,在所有数据集、变体和模型中,文本+标签的F1性能提高了0.9%到13.2%。在不同实体类型、HTML密度和构造质量的数据集上,这种性能的提高表明我们的方法是灵活和适应性强的。这些发现意味着类似的技术可能用于其他Web感知的NLP任务,包括丰富深度语言模型。
HTML tags are typically discarded in free text Named Entity Recognition from Web pages. We investigate whether these discarded tags might be used to improve NER performance. We compare Text+Tags sentences with their Text-Only equivalents, over five datasets, two free text segmentation granularities and two NER models. We find an increased F1 performance for Text+Tags of between 0.9% and 13.2% over all datasets, variants and models. This performance increase, over datasets of varying entity types, HTML density and construction quality, indicates our method is flexible and adaptable. These findings imply that a similar technique might be of use in other Web-aware NLP tasks, including the enrichment of deep language models.