A short text modeling method combining semantic and statistical information

A short text modeling method combining semantic and statistical information
复制标题

DOI:
10.1016/j.ins.2010.06.021
复制
发表时间:
2010-10
期刊:
Inf. Sci.
影响因子:
--
通讯作者:
Wenyin Liu;Xiaojun Quan;Min Feng;B. Qiu
Wenyin Liu;Xiaojun Quan;Min Feng;B. Qiu
中科院分区:
其他
文献类型:
--
作者:
Wenyin Liu;Xiaojun Quan;Min Feng;B. Qiu

文献摘要

被引文献

相似文献

提出了一种新的短文本片断集合建模方法,用于度量片断对之间的相似性。该方法同时考虑了短文本片段中的语义信息和统计信息,并由三个步骤组成。给定一组原始的短文本片段,它首先使用词汇数据库建立单词之间的初始相似度。然后,该方法迭代地计算单词相似度和短文本相似度。最后,基于词的相似度构造邻近度矩阵,将原始文本片段转换为向量。词相似度和文本聚类实验表明,本文提出的短文本建模方法提高了现有文本相关信息检索技术的性能。
A novel modeling method for a collection of short text snippets is presented in this paper to measure the similarity between pairs of snippets. The method takes account of both the semantic and statistical information within the short text snippets, and consists of three steps. Given a set of raw short text snippets, it first establishes the initial similarity between words by using a lexical database. The method then iteratively calculates both word similarity and short text similarity. Finally, a proximity matrix is constructed based on word similarity and used to convert the raw text snippets into vectors. Word similarity and text clustering experiments show that the proposed short text modeling method improves the performance of existing text-related information retrieval (IR) techniques.