WebKhoj: Indian language IR from multiple character encodings

WebKhoj: Indian language IR from multiple character encodings
复制标题

WebKhoj:来自多种字符编码的印度语言 IR

DOI:
--
复制
发表时间:
2006
期刊:
The Web Conference
影响因子:
--
通讯作者:
Vasudeva Varma
Vasudeva Varma
中科院分区:
--
文献类型:
--
作者:
Prasad Pingali;Jagadeesh Jagarlamudi;Vasudeva Varma

文献摘要

被引文献

相似文献

今天,网络搜索引擎提供了最简单的方式来获取网络上的信息。在这种情况下,超过95%的印度语内容在网络上是不可搜索的,由于网页的多种编码。这些编码大多是专有的,因此需要某种标准化,使内容通过搜索引擎访问。在本文中,我们提出了一个名为WebKhoj的搜索引擎,它能够在网络上搜索多脚本和多编码的印度语言内容。我们描述了一个语言为重点的爬虫和转码过程中涉及到实现印度语言内容的可访问性。最后,我们报告了一些实验,进行了沿着与印度语的网页内容的结果。
Today web search engines provide the easiest way to reach information on the web. In this scenario, more than 95% of Indian language content on the web is not searchable due to multiple encodings of web pages.Most of these encodings are proprietary and hence need some kind of standardization for making the content accessible via a search engine. In this paper we present a search engine called WebKhoj which is capable of searching multi-script and multi-encoded Indian language content on the web. We describe a language focused crawler and the transcoding processes involved to achieve accessibility of Indian langauge content. In the end we report some of the experiments that were conducted along with results on Indian language web content.