Combining Word Embeddings with Taxonomy Information for Multi-Label Document Classification

Combining Word Embeddings with Taxonomy Information for Multi-Label Document Classification
复制标题

将词嵌入与分类信息相结合以进行多标签文档分类

DOI:
10.1145/3342558.3345424
复制
发表时间:
2019
期刊:
ACM Symposium on Document Engineering
影响因子:
--
通讯作者:
D. Schoder
D. Schoder
中科院分区:
--
文献类型:
--
作者:
Stefan Hirschmeier;D. Schoder

文献摘要

参考文献

被引文献

相似文献

在业务上下文中,通常需要使用公司特定的分类法对文档进行分类。基于词嵌入的文本分类方法已经变得越来越流行,因为它们使词、文档和标签能够以语义上鲁棒的方式表示(作为其上下文的分布式表示),并且使文档和标签在代数向量空间中可处理。然而,这些分布式的上下文表示在用于多标签分类任务时有其缺点:两个标签的上下文越相似,它们在分类中分离就越困难。在实践中,由于训练数据不佳、训练不佳或词嵌入方法的固有局限性,我们会发现一些不可预测的区域,从而导致误报预测(通常在分类树的叶标签中)。我们贡献了一种方法来解决的问题,难以区分的领域的多标签分类任务的基础上,词嵌入,包括分类信息在预测。
In business contexts, documents often need to be classified using company-specific taxonomies. Text-classification approaches based on word embeddings have become increasingly popular as they enable words, documents, and tags to be represented in a semantically robust way (as distributed representations of their contexts) and make documents and tags processable in an algebraic vector space. However, these distributed representations of contexts have their shortcomings when used for multi-label classification tasks: the more similar the contexts of two tags, the more difficult they are to separate in classification. Intensified by poor training data, poor training, or inherent limitations of the word-embedding approach, in practice, we find areas of indistinguishability, leading to false positive predictions (typically in leaf tags of a taxonomy tree). We contribute an approach to tackle the problem of indistinguishable areas for multi-label classification tasks based on word embeddings by including taxonomy information during prediction.
Le Fort I截骨固定方法及术后骨片移位检查
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者:
藤尾正人;佐世暁;荻須宏太;土屋周平;酒井陽;日比英晴
通讯作者: 日比英晴