On the unsupervised analysis of domain-specific Chinese texts

On the unsupervised analysis of domain-specific Chinese texts
复制标题

特定领域中文文本的无监督分析

DOI:
10.1073/pnas.1516510113
复制
发表时间:
2016-05-31
影响因子:
11.1
通讯作者:
Liu,Jun S.
Liu,Jun S.
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Deng,Ke;Bol,Peter K.;Liu,Jun S.

文献摘要

被引文献

相似文献

随着数字化文本数据在公共和私人领域的日益可用性,迫切需要有效的计算工具来自动从文本中提取信息。由于中文语言与基于字母的语言最显著的不同之处在于不指定词边界,因此大多数现有的中文文本挖掘方法需要预先指定的词汇表和/或大的相关训练语料库,这在某些应用中可能不可用。我们介绍了一种无监督的方法,自上而下的词发现和分割(TopWORDS),同时发现和分割的单词和短语,从大量的非结构化的中文文本,并提出了方法,发现的话,并进行更高层次的上下文分析。TopWORDS对于挖掘在线和特定领域的文本特别有用,这些文本的基础词汇是未知的,或者感兴趣的文本与可用的训练语料库有很大差异。当TopWORDS的输出被馈送到上下文分析工具(如主题建模,单词嵌入和关联模式发现)时,结果与使用监督分割方法的输出一样好或更好。
With the growing availability of digitized text data both publicly and privately, there is a great need for effective computational tools to automatically extract information from texts. Because the Chinese language differs most significantly from alphabet-based languages in not specifying word boundaries, most existing Chinese text-mining methods require a prespecified vocabulary and/or a large relevant training corpus, which may not be available in some applications. We introduce an unsupervised method, top-down word discovery and segmentation (TopWORDS), for simultaneously discovering and segmenting words and phrases from large volumes of unstructured Chinese texts, and propose ways to order discovered words and conduct higher-level context analyses. TopWORDS is particularly useful for mining online and domain-specific texts where the underlying vocabulary is unknown or the texts of interest differ significantly from available training corpora. When outputs from TopWORDS are fed into context analysis tools such as topic modeling, word embedding, and association pattern finding, the results are as good as or better than that from using outputs of a supervised segmentation method.