A hybrid method for Chinese address segmentation

A hybrid method for Chinese address segmentation
复制标题

DOI:
10.1080/13658816.2017.1379084
复制
发表时间:
2018-01
影响因子:
5.7
通讯作者:
Lin Li;Wei Wang;B. He;Yu Zhang
Lin Li;Wei Wang;B. He;Yu Zhang
中科院分区:
地球科学2区
文献类型:
--
作者:
Lin Li;Wei Wang;B. He;Yu Zhang

文献摘要

被引文献

相似文献

中文地址切分是地理信息系统地理编码中面临的严峻挑战。以前的大多数研究都依赖于预先定义的地名词典,而没有考虑原始地址语料库包含的信息。本文提出了一种基于规则和统计方法的混合方法,用于在没有预定义地名词典的情况下进行中文地址分割。该方法使用统计方法从原始地址语料库中提取地址信息,并使用基于规则的方法对中文地址进行切分。中国,对两种典型的统计方法及其与基于规则的方法的组合与混合方法进行了比较,实验涉及深圳市约46万个地址项。实验结果表明,该方法的F值在0.8以上,优于已有的方法,从而验证了该方法的有效性。
ABSTRACT Chinese address segmentation is a serious challenge in geographic information system geocoding. Most previous studies have relied on predefined gazetteers without considering the information contained by a raw address corpus. In this paper, a hybrid method employing both rule-based and statistical methods is proposed for Chinese address segmentation without a predefined gazetteer. This approach utilizes statistical methods to extract address information from a raw address corpus and a rule-based method to segment Chinese addresses. Two typical statistical methods and their combinations with rule-based methods are compared with the hybrid method in an experiment involving approximately 460,000 address items in Shenzhen City, China. The experimental results indicate that the proposed method achieves an F-score of over 0.8, which is better than those of existing methods, thus validating the proposed method.