A Semi-supervised Approach of Extracting Attribute-Value Pairs of Chinese eBook using Conditional Random Fields

A Semi-supervised Approach of Extracting Attribute-Value Pairs of Chinese eBook using Conditional Random Fields
复制标题

DOI:
10.4156/jcit.vol8.issue1.28
复制
发表时间:
2013-01
期刊:
Journal of Convergence Information Technology
影响因子:
--
通讯作者:
Yongquan Dong;Ping Ling;Qiang Chu
Yongquan Dong;Ping Ling;Qiang Chu
中科院分区:
其他
文献类型:
--
作者:
Yongquan Dong;Ping Ling;Qiang Chu

文献摘要

相似文献

我们描述了一种方法来提取属性值对中文电子书的描述,以增加图书数据库表示为一组属性值对的每本电子书。这样的表示对于诸如需求预测、产品推荐之类的任务是有益的。目前的属性值提取方法有:基于规则的方法和基于机器学习的方法。由于大多数中文电子书描述没有统一的结构,依赖规则的方法似乎不能令人满意地执行。在本文中,我们将抽取任务描述为一个序列标记问题,并使用一个带有条件随机场(CRF)的半监督算法来解决它。抽取系统需要有限数量的标记训练样本,以减少人工准备训练样本的工作。在抽取系统中,我们使用了丰富的特征,包括文字,上下文和语义。最后,使用约束条件将提取的属性和值链接成表单对。实验结果表明,该方法对中文电子书属性值对的抽取具有良好的性能。
We describe a method to extract attribute–value pairs from Chinese eBook descriptions in order to augment book databases by representing each eBook as a set of attribute-value pairs. Such a representation is beneficial for tasks such as demand forecasting, product recommendations. Current attribute-value extraction approaches include: rule based or machine learning based. Since there is no consolidated structure for most Chinese eBook descriptions, the approach relying on rules does not seem to perform satisfactorily. In this paper, we formulate the extraction task as a sequential labeling problem and use a semi-supervised algorithm with Conditional Random Fields (CRF) to solve it. The extraction system requires a limited amount of labeled training examples to reduce the human work in preparing training examples. In the extraction system, we use rich features including literal, context and semantics. Finally, the extracted attributes and values are linked to form pairs using constraint conditions. Experimental results show that our proposed method has good performance to extract Chinese eBook attribute-value pairs.