Construction of concept network from large numbers of texts for information examination using TF-IDF and deletion of unrelated words

Construction of concept network from large numbers of texts for information examination using TF-IDF and deletion of unrelated words
复制标题

DOI:
10.1109/scis-isis.2014.7044701
复制
发表时间:
2014-12
期刊:
2014 Joint 7th International Conference on Soft Computing and Intelligent Systems (SCIS) and 15th International Symposium on Advanced Intelligent Systems (ISIS)
影响因子:
--
通讯作者:
Yuta Doen;M. Murata;Ryuta Otake;M. Tokuhisa;Qing Ma
Yuta Doen;M. Murata;Ryuta Otake;M. Tokuhisa;Qing Ma
中科院分区:
其他
文献类型:
--
作者:
Yuta Doen;M. Murata;Ryuta Otake;M. Tokuhisa;Qing Ma

文献摘要

相似文献

我们提出了一种新的方法来构建一个网络,该网络描述电子文本中与某个关键字相关的事物之间的关系信息。该方法具有两个特点(TF-IDF和删除无关词)。我们使用了一种基于词频-逆文档频率(TF-IDF)的方法来提取相关词。使用TF-IDF,我们只提取重要的单词。我们使用TF-IDF作为网络中一条边的权重。我们还删除网络中不相关的单词。当扩展网络和添加单词时,可能会添加不相关的单词。该系统使用主题受限和主题相关两种方法来删除这些不相关的词。实验证明,本文提出的基于TF-IDF的关联词提取方法比基于条件概率的关联词提取方法获得了更好的结果。我们还进行了实验,验证了删除不相关词的有效性。我们发现,主题限制方法可以删除大部分不相关的词,并从原始网络中保留大约0.8%的相关词。主题相关方法可以删除一些不相关的词,并保留原网络中的大多数相关词。
We propose new methods to construct a network that describes information about the relations of things that are related to a certain keyword from electronic texts. The proposed method has two characteristics (TF-IDF and deletion of unrelated words). We extract related words using a term frequency-inverse document frequency (TF-IDF)-based method. Using TF-IDF, we extract only important words. We use TF-IDF as a weight for an edge in a network. We also delete unrelated words in the network. When expanding a network and adding words, unrelated words are likely to be added. The proposed system deletes such unrelated words using two methods, the topic-restricted and topic-related methods. We have experimentally confirmed that the proposed TF-IDF-based related word extraction method obtains better results than a method that uses conditional probabilities to extract related words. We also conducted experiments to verify the effectiveness of deleting unrelated words. We found that the topic-restricted method could delete most unrelated words and maintain approximately 0.8 of the related words from the original network. The topic-related method can delete some unrelated words and maintain most related words from the original network.