课题基金 / 基金详情

III: SpatioTextual Extraction of Document on the Web for Digital Government Applications

III: SpatioTextual Extraction of Document on the Web for Digital Government Applications
III:用于数字政府应用的网络文档的空间文本提取
批准号:
0713501
负责人:
Hanan Samet
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-09-01 至 2011-08-31

项目摘要

项目成果

Hanan Samet的其他基金

相关文献

中文摘要
翻译
当今的搜索技术主要是由搜索引擎,如谷歌提供的搜索引擎,对于给定的查询字符串s,一组文档D在一种算法的帮助下被检索出来,该算法根据有多少其他文档链接到它,对D的元素进行排名。本研究将研究支持地理位置检索的搜索引擎的开发中涉及的问题,以及其在涉及数字政府应用的环境中的部署,在该环境中还期望基于空间邻近度来检索文档。知识价值:(1)识别文档中的地理参考是一个具有挑战性的问题,对于高级搜索应用程序是必要的。这是特别真实的情况下,用户想要浏览大量的文件,不一定在网络上,并探索和发现空间关系,无论是在同一个文件或文件的集合。(2)越来越多的用户正在寻找包含空间上接近的内容的文档。因此传统的按链接结构进行排名的方法是不合适的。确定文件的地理焦点是一项困难的任务,但在处理隐藏网络上的文件等应用程序中是必要的,隐藏网络是一组文件,通常是专有的,供组织内部使用,通常在互联网上无法获得。这意味着这些文件的链接很少,因此流行的互联网搜索策略不适用。(3)将文档的空间内容视为一等公民,在这个意义上,无论查询是否具有空间成分,都为检索到的每个文档报告地理范围,由于需要解决与混叠相关的问题,(意识到“洛杉矶”和“洛杉矶”是一样的)和歧义(对“伦敦”的不同解释)。(4)为涉及文本和空间组件的查询开发查询优化和执行策略。(5)开发有效的技术来测量空间相似性而不是接近性,以及测量空间和文本相似性组合的技术。这包括天际线运营商的适应。广泛影响:根据空间接近度检索文件的能力有助于改善搜索体验,并将产生更相关的结果。开发的工具还将扩大搜索引擎的范围,从仅限于互联网上的文档到隐藏网络上的文档。通过与赠款的数字政府合作伙伴合作,在政府网站上部署这些工具,其效果是使公民能够了解他们的政府正在做什么,从而使公民更加知情。
英文摘要
Search technology today is dominated by search engines such as the one provided by Google where for a given query string s , a set D of documents is retrieved with the aid of an algorithm that ranks the elements of D on the basis of how many other documents link to it. This research will investigate the issues involved in the development of a search engine that supports geographic location retrieval, and its deployment in a setting involving digital government applications where it is also desirable to retrieve documents on the basis of spatial proximity. Intellectual Merit: (1) Identifying geographic references in documents is a challenging issue and is necessary for advanced search applications. This is especially true in the case of users who want to browse large collections of documents that are not necessarily on the web and to explore and discover spatial relationships either in the same document or in a collection of documents. (2) Increasingly, users are looking for documents that contain spatially proximate content. Thus the traditional method of ranking by the link structure of the web is not appropriate. Determining the geographic focus of a document is a difficult task but is necessary in applications such as those dealing with documents on the hidden web, which is a set of documents, usually proprietary, that is for internal use of an organization and is often not available on the Internet. This means that there are few, if any links to these documents, and thus popular internet search strategies are not applicable. (3) Treating spatial content of documents as a first-class citizen, in the sense that a geographic scope is reported for each document that is retrieved regardless of whether the query has a spatial component, is difficult given the need to resolve issues related to aliasing (realizing that ''''Los Angeles'''' and ''''LA'''' are the same) and ambiguity (different interpretations for ''''London''''). (4) Developing query optimization and execution strategies for queries that involve both a textual and spatial component. (5) Developing effective techniques for measuring spatial similarity other than proximity, as well as techniques for measuring combinations of spatial and textual similarity. This includes the adaptation of the skyline operator. Broad Impacts: The ability to retrieve documents on the basis of spatial proximity makes for a better search experience and will lead to more relevant results. The tools to be developed will also extend the reach of search engines from being restricted to documents on the internet to documents that reside on the hidden web. The deployment of these tools in government web sites via collaboration with the grant''s digital government partners has the effect of empowering citizens to find out what their government is doing, thereby leading to a more informed citizenry.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: Small: Trajectory Computing
EAGER: NewsStand CoronaViz: A Map Query Interface for Tracking the Spread of COVID-19
III: Small: Using Location for Retrieving Text and Images in News And Social Media Posts
I-Corps: RoadsInDB: Customer Discovery in the Logistics, Delivery, Ride Sharing, Location-based Services and Analytics Verticals