Construction of concept network from large numbers of texts for information examination using TF-IDF and deletion of unrelated words
Construction of concept network from large numbers of texts for information examination using TF-IDF and deletion of unrelated words
复制标题
DOI:
10.1109/scis-isis.2014.7044701
复制
发表时间:
2014-12
期刊:
影响因子:
--
通讯作者:
Yuta Doen;M. Murata;Ryuta Otake;M. Tokuhisa;Qing Ma
中科院分区:
文献类型:
--
作者:
Yuta Doen;M. Murata;Ryuta Otake;M. Tokuhisa;Qing Ma
We propose new methods to construct a network that describes information about the relations of things that are related to a certain keyword from electronic texts. The proposed method has two characteristics (TF-IDF and deletion of unrelated words). We extract related words using a term frequency-inverse document frequency (TF-IDF)-based method. Using TF-IDF, we extract only important words. We use TF-IDF as a weight for an edge in a network. We also delete unrelated words in the network. When expanding a network and adding words, unrelated words are likely to be added. The proposed system deletes such unrelated words using two methods, the topic-restricted and topic-related methods. We have experimentally confirmed that the proposed TF-IDF-based related word extraction method obtains better results than a method that uses conditional probabilities to extract related words. We also conducted experiments to verify the effectiveness of deleting unrelated words. We found that the topic-restricted method could delete most unrelated words and maintain approximately 0.8 of the related words from the original network. The topic-related method can delete some unrelated words and maintain most related words from the original network.