Mining the Frequent Patterns of Named Entities for Long Document Classification

Mining the Frequent Patterns of Named Entities for Long Document Classification
复制标题

DOI:
10.3390/app12052544
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Wenjun Ke
Wenjun Ke
中科院分区:
--
文献类型:
--
作者:
Bohan Wang;Rui Qi;Jinhua Gao;Jianwei Zhang;Xiaoguang Yuan;Wenjun Ke

文献摘要

相似文献

Nowadays, a large amount of information is stored as text, and numerous text mining techniques have been developed for various applications, such as event detection, news topic classification, public opinion detection, and sentiment analysis. Although significant progress has been achieved for short text classification, document-level text classification requires further exploration. Long documents always contain irrelevant noisy information that shelters the prominence of indicative features, limiting the interpretability of classification results. To alleviate this problem, a model called MIPELD (mining the frequent pattern of a named entity for long document classification) for long document classification is demonstrated, which mines the frequent patterns of named entities as features. Discovered patterns allow semantic generalization among documents and provide clues for verifying the results. Experiments on several datasets resulted in good accuracy and marco-F1 values, meeting the requirements for practical application. Further analysis validated the effectiveness of MIPELD in mining interpretable information in text classification.