课题基金 / 基金详情

III: SpatioTextual Extraction of Document on the Web for Digital Government Applications

III: SpatioTextual Extraction of Document on the Web for Digital Government Applications
III:用于数字政府应用的网络文档的空间文本提取
批准号:
0713501
负责人:
Hanan Samet
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-09-01 至 2011-08-31

项目摘要

项目成果

Hanan Samet的其他基金

相关文献

中文摘要
翻译
今天的搜索技术是由谷歌提供的搜索引擎主导的,在谷歌中,对于给定的查询字符串s,在一种算法的帮助下检索一组D的文档,该算法根据链接到D的其他文档的数量对D的元素进行排名。本研究将探讨支持地理位置检索的搜索引擎的开发所涉及的问题,以及它在涉及数字政府应用程序的设置中的部署,其中也需要基于空间接近度检索文档。智力优势:(1)识别文档中的地理参考是一个具有挑战性的问题,对于高级搜索应用程序是必要的。如果用户想要浏览不一定在web上的大型文档集合,并且想要探索和发现同一文档或文档集合中的空间关系,则尤其如此。(2)越来越多的用户正在寻找包含空间近似内容的文档。因此,传统的通过网页链接结构进行排名的方法是不合适的。确定文档的地理焦点是一项困难的任务,但在诸如处理隐藏web上的文档的应用程序中是必要的,隐藏web是一组文档,通常是专有的,供组织内部使用,通常在Internet上不可用。这意味着这些文档的链接很少(如果有的话),因此流行的互联网搜索策略不适用。(3)将文档的空间内容视为一级公民,即无论查询是否具有空间组件,都会为检索到的每个文档报告一个地理范围,考虑到需要解决与混联(认识到‘’‘’Los Angeles‘‘‘’和’’‘’LA‘‘‘’是相同的)和歧义(对’’‘’London''‘’的不同解释)相关的问题,很难将文档的空间内容视为一级公民。(4)为包含文本和空间组件的查询开发查询优化和执行策略。(5)开发除接近性以外的有效的空间相似性测量技术,以及空间和文本相似性组合的测量技术。这包括天际线操作符的改编。广泛的影响:基于空间接近度检索文档的能力可以提供更好的搜索体验,并将导致更相关的结果。这些有待开发的工具还将把搜索引擎的范围从仅限于互联网上的文档扩展到隐藏网络上的文档。通过与赠款的数字政府合作伙伴合作,在政府网站上部署这些工具,其效果是使公民能够了解他们的政府正在做什么,从而使公民更加知情。
英文摘要
Search technology today is dominated by search engines such as the one provided by Google where for a given query string s , a set D of documents is retrieved with the aid of an algorithm that ranks the elements of D on the basis of how many other documents link to it. This research will investigate the issues involved in the development of a search engine that supports geographic location retrieval, and its deployment in a setting involving digital government applications where it is also desirable to retrieve documents on the basis of spatial proximity. Intellectual Merit: (1) Identifying geographic references in documents is a challenging issue and is necessary for advanced search applications. This is especially true in the case of users who want to browse large collections of documents that are not necessarily on the web and to explore and discover spatial relationships either in the same document or in a collection of documents. (2) Increasingly, users are looking for documents that contain spatially proximate content. Thus the traditional method of ranking by the link structure of the web is not appropriate. Determining the geographic focus of a document is a difficult task but is necessary in applications such as those dealing with documents on the hidden web, which is a set of documents, usually proprietary, that is for internal use of an organization and is often not available on the Internet. This means that there are few, if any links to these documents, and thus popular internet search strategies are not applicable. (3) Treating spatial content of documents as a first-class citizen, in the sense that a geographic scope is reported for each document that is retrieved regardless of whether the query has a spatial component, is difficult given the need to resolve issues related to aliasing (realizing that ''''Los Angeles'''' and ''''LA'''' are the same) and ambiguity (different interpretations for ''''London''''). (4) Developing query optimization and execution strategies for queries that involve both a textual and spatial component. (5) Developing effective techniques for measuring spatial similarity other than proximity, as well as techniques for measuring combinations of spatial and textual similarity. This includes the adaptation of the skyline operator. Broad Impacts: The ability to retrieve documents on the basis of spatial proximity makes for a better search experience and will lead to more relevant results. The tools to be developed will also extend the reach of search engines from being restricted to documents on the internet to documents that reside on the hidden web. The deployment of these tools in government web sites via collaboration with the grant''s digital government partners has the effect of empowering citizens to find out what their government is doing, thereby leading to a more informed citizenry.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: Small: Trajectory Computing
EAGER: NewsStand CoronaViz: A Map Query Interface for Tracking the Spread of COVID-19
III: Small: Using Location for Retrieving Text and Images in News And Social Media Posts
I-Corps: RoadsInDB: Customer Discovery in the Logistics, Delivery, Ride Sharing, Location-based Services and Analytics Verticals