Thomson Legal and Regulatory at NTCIR-3: Japanese, Chinese and English Retrieval Experiments

Thomson Legal and Regulatory at NTCIR-3: Japanese, Chinese and English Retrieval Experiments
复制标题

Thomson Legal and Regulatory at NTCIR-3:日语、中文和英语检索实验

DOI:
--
复制
发表时间:
2002
期刊:
--
影响因子:
--
通讯作者:
P. Jackson
P. Jackson
中科院分区:
--
文献类型:
--
作者:
Isabelle Moulinier;Hugo Molina;P. Jackson

文献摘要

被引文献

相似文献

Thomson法律的和监管部门参与了NTCIR-3研讨会的CLIR任务。我们提交了正式运行的单语检索在日本和中国,从英语到日语的双语检索。我们主要关注的是日语检索。我们比较了基于单词和基于字符的索引,以及使用字符和字符二元组的查询公式。我们的研究结果表明,基于词和基于双字母的检索表现出类似的性能,大多数查询制定的方法,而他们优于基于字符的检索。在中文检索方面,我们比较了单字符检索和双字符检索。我们还引入了一个结构化查询来利用两者。我们的结果与以前的工作一致,其中字符二元组被证明比单个字符具有更好的性能。结构化查询方法很有前途,但需要更多的分析。在我们的双语运行中,查询使用机器可读词典进行翻译。翻译后的术语被重新分段以匹配索引单位。到目前为止,我们的结果是不确定的,因为我们经历了意想不到的查询公式化问题,特别是在我们的基于单词的方法。
Thomson Legal and Regulatory participated in the CLIR task of the NTCIR-3 workshop. We submitted formal runs for monolingual retrieval in Japanese and Chinese, and for bilingual retrieval from English to Japanese. Our main focus was in Japanese retrieval. We compared word-based and character-based index- ing, as well as query formulation using characters and character bigrams. Our results show that word- based and bigram-based retrieval show similar perfor- mance for most query formulation approaches, while they outperform character-based retrieval. For Chi- nese retrieval, we compared using single characters with using character bigrams. We also introduced a structured query to leverage both. Our results are con- sistent with previous work, where character bigrams were shown to have better performance than single characters. The structured query approach is promis- ing, but requires more analysis. In our bilingual runs, queries were translated using a machine-readable dic- tionary. Translated terms were resegmented to match indexing units. Our results, so far, are inconclusive, as we experienced unexpected query formulation issues especially in our word-based approach.