Gathering Web Pages of Entities with High Precision

Gathering Web Pages of Entities with High Precision
复制标题

DOI:
--
复制
发表时间:
2014-11
期刊:
J. Web Eng.
影响因子:
--
通讯作者:
Byung-Won On;Muhammad Omar;G. Choi;Junbeom Kwon
Byung-Won On;Muhammad Omar;G. Choi;Junbeom Kwon
中科院分区:
其他
文献类型:
--
作者:
Byung-Won On;Muhammad Omar;G. Choi;Junbeom Kwon

文献摘要

被引文献

相似文献

像Yahoo这样的搜索引擎在网页上搜索实体,例如特定的人,地点或事物。根据查询关键字的粒度和搜索引擎的性能,检索到的网页可能数量非常大,具有许多不相关的网页,并且也可能不按正确的顺序。由于检索到的网页数量庞大,人工判断每个网页的相关性是不可行的。另一个挑战是开发由搜索引擎提供的搜索结果的独立于语言的相关性分类。为了提高搜索引擎的质量,希望自动评估搜索引擎的结果,并决定检索到的网页与用户查询和预期实体的相关性,查询是所有关于。这种改进的一个步骤是通过理解用户的需求来修剪不相关的网页,以便发现特定领域中的实体的知识。我们提出了一种新的方法来提高搜索引擎的精度,该方法与语言无关,并且不受搜索引擎查询日志和用户点击数据的影响(最近被广泛使用)。我们设计语言无关的新功能,建立支持向量机相关性分类模型,使用它可以自动分类是否由搜索引擎检索到的网页是相关或不需要的实体。
A search engine like Yahoo looks for entities such as specific people, places, or things on web pages with search queries. Depending on the granularity of query keywords and performance of a search engine, the retrieved web pages may be in very large number having lots of irrelevant web pages and may be also not in proper order. It's infeasible to manually decide the relevance of each web page due to the large number of retrieved web pages. Another challenge is to develop a language independent relevance classification of search results provided by a search engine. To improve the quality of a search engine it is desirable to automatically evaluate the results of a search engine and decide the relevance of retrieved web pages with the user query and the intended entity, the query is all about. A step towards this improvement is to prune irrelevant web pages out by understanding the needs of a user in order to discover knowledge of entities in a particular domain. We propose a novel method to improve the precision of a search engine which is language independent and also free from search engine query logs and user clicks through data (widely used in recent times). We devise language independent novel features to build support vector machine relevance classification model using which we can automatically classify whether a web page retrieved by a search engine is relevant or not to the desired entity.