Conservative Likelihood Ratio Estimator for Infrequent Data Slightly above a Frequency Threshold

Conservative Likelihood Ratio Estimator for Infrequent Data Slightly above a Frequency Threshold
复制标题

DOI:
10.1109/icaicta56449.2022.9932917
复制
发表时间:
2022-09
期刊:
2022 9th International Conference on Advanced Informatics: Concepts, Theory and Applications (ICAICTA)
影响因子:
--
通讯作者:
Masato Kikuchi;Yuhi Kusakabe;Tadachika Ozono
Masato Kikuchi;Yuhi Kusakabe;Tadachika Ozono
中科院分区:
其他
文献类型:
--
作者:
Masato Kikuchi;Yuhi Kusakabe;Tadachika Ozono

文献摘要

相似文献

使用观察到的事件频率的朴素似然比(LR)估计可能高估不频繁数据的LR。避免此问题的一种方法是使用频率阈值,并将低于阈值的频率的估计值设置为零。这种方法省去了一些估计的计算,从而使使用LRS的实际任务更加有效。然而,它仍然高估了接近阈值的低频LRs。这项研究为低频提出了一个保守的估计值,略高于阈值。我们的实验使用LRS来预测语料库中命名实体的出现上下文。实验结果表明,该估计器在保持上下文预测效率的同时,提高了预测精度。
A naive likelihood ratio (LR) estimation using the observed frequencies of events can overestimate LRs for infrequent data. One approach to avoid this problem is to use a frequency threshold and set the estimates to zero for frequencies below the threshold. This approach eliminates the computation of some estimates, thereby making practical tasks using LRs more efficient. However, it still overestimates LRs for low frequencies near the threshold. This study proposes a conservative estimator for low frequencies, slightly above the threshold. Our experiment used LRs to predict the occurrence contexts of named entities from a corpus. The experimental results demonstrate that our estimator improves the prediction accuracy while maintaining efficiency in the context prediction task.