Reducing Semantic Drift with Bagging and Distributional Similarity

Reducing Semantic Drift with Bagging and Distributional Similarity
复制标题

DOI:
10.3115/1687878.1687935
复制
发表时间:
2009-08
期刊:
--
影响因子:
--
通讯作者:
Tara McIntosh;J. Curran
Tara McIntosh;J. Curran
中科院分区:
其他
文献类型:
--
作者:
Tara McIntosh;J. Curran

文献摘要

被引文献

相似文献

迭代自举算法通常使用一组手工挑选的种子进行比较。然而,我们证明,性能变化很大,取决于这些种子,和有利的种子一个算法可以执行非常差,与其他人,比较不可靠。我们利用这种广泛的变化与装袋,采样自动提取的种子,以减少语义漂移。然而,语义漂移仍然发生在以后的迭代中。我们提出了一个集成的分布式相似性过滤器,以识别和审查潜在的语义漂移,确保超过10%的精度时,提取大型语义词典。
Iterative bootstrapping algorithms are typically compared using a single set of hand-picked seeds. However, we demonstrate that performance varies greatly depending on these seeds, and favourable seeds for one algorithm can perform very poorly with others, making comparisons unreliable. We exploit this wide variation with bagging, sampling from automatically extracted seeds to reduce semantic drift. However, semantic drift still occurs in later iterations. We propose an integrated distributional similarity filter to identify and censor potential semantic drifts, ensuring over 10% higher precision when extracting large semantic lexicons.