A Study of Analogical Density in Various Corpora at Various Granularity

A Study of Analogical Density in Various Corpora at Various Granularity
复制标题

DOI:
10.3390/info12080314
复制
发表时间:
2021-08
期刊:
Inf.
影响因子:
--
通讯作者:
Rashel Fam;Y. Lepage
Rashel Fam;Y. Lepage
中科院分区:
其他
文献类型:
--
作者:
Rashel Fam;Y. Lepage

文献摘要

相似文献

本文考察了计算语篇中句间类比数量的理论问题。在此基础上,我们测量了文本的类比密度。我们关注的是句子层面的类比,基于形式层面而不是语义层面。实验进行了两个不同的语料库,在六个欧洲语言已知具有不同程度的形态丰富。语料库使用几种标记化方案进行标记化:字符,子词和词。对于子词标记化方案,我们采用两种流行的子词模型:单字语言模型和字节对编码。结果表明,类词比越高的语料,其类比密度越高。我们还观察到,根据频率掩蔽标记有助于增加类比密度。至于标记化方案,结果表明,类比密度从字符到单词降低。然而,当令牌基于其频率被屏蔽时,这是不正确的。我们发现,标记的句子使用子词模型和掩蔽最不频繁的令牌增加类比密度。
In this paper, we inspect the theoretical problem of counting the number of analogies between sentences contained in a text. Based on this, we measure the analogical density of the text. We focus on analogy at the sentence level, based on the level of form rather than on the level of semantics. Experiments are carried on two different corpora in six European languages known to have various levels of morphological richness. Corpora are tokenised using several tokenisation schemes: character, sub-word and word. For the sub-word tokenisation scheme, we employ two popular sub-word models: unigram language model and byte-pair-encoding. The results show that the corpus with a higher Type-Token Ratio tends to have higher analogical density. We also observe that masking the tokens based on their frequency helps to increase the analogical density. As for the tokenisation scheme, the results show that analogical density decreases from the character to word. However, this is not true when tokens are masked based on their frequencies. We find that tokenising the sentences using sub-word models and masking the least frequent tokens increase analogical density.