Challenges in Developing Persian Corpora from Online Resources

Challenges in Developing Persian Corpora from Online Resources
复制标题

利用在线资源开发波斯语语料库的挑战

DOI:
--
复制
发表时间:
2009
期刊:
International Conference on Asian Language Processing
影响因子:
--
通讯作者:
S. Momtazi
S. Momtazi
中科院分区:
--
文献类型:
--
作者:
Masood Ghayoomi;S. Momtazi

文献摘要

被引文献

相似文献

波斯语是印欧语系的一种语言,其文字借用了闪米特语系的阿拉伯语。由于波斯语和阿拉伯语的文字非常相似,当我们想要处理电子文本时就会出现问题。在本文中,一些共同面临的问题,在实验中发展的语料库波斯语在线材料进行了讨论。问题的根源是波斯文字本身;与阿拉伯文字的混合;波斯正字法;打字员的打字风格;以及操作系统中波斯代码页与阿拉伯代码页的混合。
Persian is one of the Indo-European languages which has borrowed its script from Arabic, a member of Semitic language family. Since Persian and Arabic scripts are so similar, problems arise when we want to process an electronic text. In this paper, some of the common problems faced experimentally in developing a corpus for Persian from on-line materials are discussed. The sources of the problems are the Persian script itself; mixture with the Arabic script; Persian orthography; the typists’ typing styles; and mixing Persian code pages with Arabic code pages in operating systems.