RIA: A Testbed for the Application of Corpus Linguistics to Information Retrieval
RIA: A Testbed for the Application of Corpus Linguistics to Information Retrieval
批准号:
9409263
负责人:
Susan Gauch
金额:
$10.49万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1994
资助国家:
美国
项目状态:
已结题
起止时间:
1994-08-15 至 1998-08-31
中文摘要
快速增长的存储介质容量和不断扩展的互连性预示着信息时代的到来。 不幸的是,获取在线信息仍然是一门不精确的科学。 虽然可以找到有价值的信息,但通常也会检索到许多不相关的文档,并且会遗漏许多相关的文档。 用户查询和文档内容之间的术语不匹配是检索失败的原因之一。 使用相关词扩展用户查询可以提高搜索性能,但识别相关词的问题仍然存在。 本研究使用语料库语言学技术,自动发现词的相似性直接从一个未标记的文本数据库的内容,并将该信息在信息检索系统。 这些相似性是根据单词出现的上下文计算的。 使用这些相似性,用户查询自动扩展,导致概念检索,而不是要求查询和文档之间的精确单词匹配。 评估了使用不同算法计算相似度的效果以及扩展不同查询词集的效果。 此外,检索引擎的搜索性能作为一种基于任务的方法,用于比较使用不同语料库语言学技术计算的词-词相似度的质量。
英文摘要
Rapidly increasing storage media capabilities and spreading interconnectivity have heralded the arrival of the information age. Unfortunately, accessing online information remains an inexact science. While valuable information can be found, typically many irrelevant documents are also retrieved and many relevant ones are missed. Terminology mismatches between the user's query and document contents are one cause of retrieval failures. Expanding a user's query with related words can improve search performance, but the problem of identifying related words remains. This research uses corpus linguistics techniques to automatically discover word similarities directly from the contents of an untagged textual database and to incorporate that information in an information retrieval system. These similarities are calculated based on the contexts in which the words appear. Using these similarities, user queries are automatically expanded, resulting in conceptual retrieval rather than requiring exact word matches between queries and documents. The effects of using different algorithms to calculate the similarities and the effects of expanding different sets of query words is evaluated. In addition, the search performance of the retrieval engine serves as a task-based method for comparing the quality of word-word similarities calculated using different corpus linguistics techniques.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: CI-ADDO-EN: Semantic CiteseerX
-
批准号:0958123
-
项目类别:Continuing Grant
-
资助金额:$26.25万
-
财政年份:2010
-
负责人:Susan Gauch
-
依托单位:
III: EAGER: Mapping Three-Dimensional Virtual Worlds
-
批准号:1050801
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2010
-
负责人:Susan Gauch
-
依托单位:
Supporting Students Attending the 2008 Adaptive Hypermedia Doctoral Consortium
-
批准号:0824712
-
项目类别:Standard Grant
-
资助金额:$1.46万
-
财政年份:2008
-
负责人:Susan Gauch
-
依托单位:
CRI: Collaborative: Next Generation CiteSeer
-
批准号:0800562
-
项目类别:Continuing Grant
-
资助金额:$8.96万
-
财政年份:2007
-
负责人:Susan Gauch
-
依托单位:
CRI: Collaborative: Next Generation CiteSeer
-
批准号:0454121
-
项目类别:Continuing Grant
-
资助金额:$23.0万
-
财政年份:2005
-
负责人:Susan Gauch
-
依托单位:
Biodiversity and Ecosystem Informatics (BDEI): Biodiversity Information Organization Using Taxonomy [BIOT]
-
批准号:0131835
-
项目类别:Standard Grant
-
资助金额:$9.98万
-
财政年份:2002
-
负责人:Susan Gauch
-
依托单位:
CAREER/EPSCoR: Cooperative Agents for Conceptual Search and Browsing of World Wide Web Resources
-
批准号:9703307
-
项目类别:Continuing Grant
-
资助金额:$20.0万
-
财政年份:1997
-
负责人:Susan Gauch
-
依托单位:
海外基金