RIA: A Testbed for the Application of Corpus Linguistics to Information Retrieval
RIA: A Testbed for the Application of Corpus Linguistics to Information Retrieval
批准号:
9409263
负责人:
Susan Gauch
金额:
$10.49万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1994
资助国家:
美国
项目状态:
已结题
起止时间:
1994-08-15 至 1998-08-31
中文摘要
快速增长的存储媒体能力和广泛的互联性预示着信息时代的到来。不幸的是,获取在线信息仍然是一门不精确的科学。虽然可以找到有价值的信息,但通常也会检索到许多不相关的文档,并且会丢失许多相关的文档。用户查询和文档内容之间的术语不匹配是导致检索失败的一个原因。用相关单词扩展用户的查询可以提高搜索性能,但是识别相关单词的问题仍然存在。本研究使用语料库语言学技术直接从未标记文本数据库的内容中自动发现单词相似度,并将该信息合并到信息检索系统中。这些相似度是根据单词出现的上下文计算出来的。使用这些相似性,用户查询可以自动扩展,从而实现概念检索,而不需要查询和文档之间的精确单词匹配。评估了使用不同算法计算相似度的效果以及扩展不同查询词集的效果。此外,检索引擎的搜索性能作为一种基于任务的方法,用于比较使用不同语料库语言学技术计算的词-词相似度的质量。
英文摘要
Rapidly increasing storage media capabilities and spreading interconnectivity have heralded the arrival of the information age. Unfortunately, accessing online information remains an inexact science. While valuable information can be found, typically many irrelevant documents are also retrieved and many relevant ones are missed. Terminology mismatches between the user's query and document contents are one cause of retrieval failures. Expanding a user's query with related words can improve search performance, but the problem of identifying related words remains. This research uses corpus linguistics techniques to automatically discover word similarities directly from the contents of an untagged textual database and to incorporate that information in an information retrieval system. These similarities are calculated based on the contexts in which the words appear. Using these similarities, user queries are automatically expanded, resulting in conceptual retrieval rather than requiring exact word matches between queries and documents. The effects of using different algorithms to calculate the similarities and the effects of expanding different sets of query words is evaluated. In addition, the search performance of the retrieval engine serves as a task-based method for comparing the quality of word-word similarities calculated using different corpus linguistics techniques.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: CI-ADDO-EN: Semantic CiteseerX
-
批准号:0958123
-
项目类别:Continuing Grant
-
资助金额:$26.25万
-
财政年份:2010
-
负责人:Susan Gauch
-
依托单位:
III: EAGER: Mapping Three-Dimensional Virtual Worlds
-
批准号:1050801
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2010
-
负责人:Susan Gauch
-
依托单位:
Supporting Students Attending the 2008 Adaptive Hypermedia Doctoral Consortium
-
批准号:0824712
-
项目类别:Standard Grant
-
资助金额:$1.46万
-
财政年份:2008
-
负责人:Susan Gauch
-
依托单位:
CRI: Collaborative: Next Generation CiteSeer
-
批准号:0800562
-
项目类别:Continuing Grant
-
资助金额:$8.96万
-
财政年份:2007
-
负责人:Susan Gauch
-
依托单位:
CRI: Collaborative: Next Generation CiteSeer
-
批准号:0454121
-
项目类别:Continuing Grant
-
资助金额:$23.0万
-
财政年份:2005
-
负责人:Susan Gauch
-
依托单位:
Biodiversity and Ecosystem Informatics (BDEI): Biodiversity Information Organization Using Taxonomy [BIOT]
-
批准号:0131835
-
项目类别:Standard Grant
-
资助金额:$9.98万
-
财政年份:2002
-
负责人:Susan Gauch
-
依托单位:
CAREER/EPSCoR: Cooperative Agents for Conceptual Search and Browsing of World Wide Web Resources
-
批准号:9703307
-
项目类别:Continuing Grant
-
资助金额:$20.0万
-
财政年份:1997
-
负责人:Susan Gauch
-
依托单位:
海外基金