Partially Distribution-Free Learning of Regular Languages from Positive Samples

Partially Distribution-Free Learning of Regular Languages from Positive Samples
复制标题

从正样本中进行部分无分布的正则语言学习

DOI:
10.3115/1220355.1220368
复制
发表时间:
2004
期刊:
International Conference on Computational Linguistics
影响因子:
--
通讯作者:
F. Thollard
F. Thollard
中科院分区:
--
文献类型:
--
作者:
Alexander Clark;F. Thollard

文献摘要

被引文献

相似文献

尽管常规语言存在缺陷,但它们在今天的自然语言处理中得到了广泛的使用。高效的算法可以可靠地学习这些语言,并且在现实应用中必须只使用正样本,这是必要的。在传统的免费分发标准下,这些语言是不可学习的。我们认为一个合适的学习框架是PAC学习,其中分布被约束为由一类支持等于目标概念的随机自动机生成。我们讨论了这与其他学习范式之间的关系。然后,我们给出了一个简单的正则语言学习算法,并给出了一个完备的证明,证明了它是按照这个部分分布自由准则学习的。
Regular languages are widely used in NLP today in spite of their shortcomings. Efficient algorithms that can reliably learn these languages, and which must in realistic applications only use positive samples, are necessary. These languages are not learnable under traditional distribution free criteria. We claim that an appropriate learning framework is PAC learning where the distributions are constrained to be generated by a class of stochastic automata with support equal to the target concept. We discuss how this is related to other learning paradigms. We then present a simple learning algorithm for regular languages, and a self-contained proof that it learns according to this partially distribution free criterion.