WebKhoj: Indian language IR from multiple character encodings
WebKhoj: Indian language IR from multiple character encodings
复制标题
WebKhoj:来自多种字符编码的印度语言 IR
DOI:
--
复制
发表时间:
2006
期刊:
影响因子:
--
通讯作者:
Vasudeva Varma
中科院分区:
文献类型:
--
作者:
Prasad Pingali;Jagadeesh Jagarlamudi;Vasudeva Varma
Today web search engines provide the easiest way to reach information on the web. In this scenario, more than 95% of Indian language content on the web is not searchable due to multiple encodings of web pages.Most of these encodings are proprietary and hence need some kind of standardization for making the content accessible via a search engine. In this paper we present a search engine called WebKhoj which is capable of searching multi-script and multi-encoded Indian language content on the web. We describe a language focused crawler and the transcoding processes involved to achieve accessibility of Indian langauge content. In the end we report some of the experiments that were conducted along with results on Indian language web content.