Semi-supervised Bibliographic Element Segmentation with Latent Permutations

Semi-supervised Bibliographic Element Segmentation with Latent Permutations
复制标题

DOI:
10.1007/978-3-642-24826-9_11
复制
发表时间:
2011-10
期刊:
--
影响因子:
--
通讯作者:
Tomonari Masada;A. Takasu;Yuichiro Shibata;K. Oguri
Tomonari Masada;A. Takasu;Yuichiro Shibata;K. Oguri
中科院分区:
其他
文献类型:
--
作者:
Tomonari Masada;A. Takasu;Yuichiro Shibata;K. Oguri

文献摘要

相似文献

提出了一种半监督式书目元素分割方法。我们的输入数据是一组大规模的书目参考文献,每个参考文献都是一个未分割的词标记序列。我们的问题是将每个参考文献划分为书目元素,例如作者、标题、期刊、页数等。我们使用类似于lda的主题模型来解决这个问题,方法是将每个词标记分配给一个主题,以便分配给同一主题的词标记引用相同的书目元素。主题赋值应满足连续性约束,即赋值给同一主题的词标记必须是连续的约束。因此,我们在之前的工作[8]中基于Chen等人设计的主题模型[3]提出了一个主题模型。该模型对LDA进行了扩展,实现了满足邻接约束的无监督主题分配。本文的主要贡献是为我们提出的模型提出了半监督学习。我们假设最多有三分之一的单词标记已经被标记。此外,我们假设有百分之几的标签可能是不正确的。实验表明,我们的半监督学习大大提高了无监督学习,实现了90%以上的分割准确率。
This paper proposes a semi-supervised bibliographic element segmentation. Our input data is a large scale set of bibliographic references each given as an unsegmented sequence of word tokens. Our problem is to segment each reference into bibliographic elements, e.g. authors, title, journal, pages, etc. We solve this problem with an LDA-like topic model by assigning each word token to a topic so that the word tokens assigned to the same topic refer to the same bibliographic element. Topic assignments should satisfy contiguity constraint, i.e., the constraint that the word tokens assigned to the same topic should be contiguous. Therefore, we proposed a topic model in our preceding work [8] based on the topic model devised by Chen et al. [3]. Our model extends LDA and realizes unsupervised topic assignments satisfying contiguity constraint. The main contribution of this paper is the proposal of asemi-supervisedlearning for our proposed model. We assume that at most one third of word tokens are already labeled. In addition, we assume that a few percent of the labels may be incorrect. The experiment showed that our semi-supervised learning improved the unsupervised learning by a large margin and achieved an over 90% segmentation accuracy.