Accounting for Sentence Position and Legal Domain Sentence Embedding in Learning to Classify Case Sentences

Accounting for Sentence Position and Legal Domain Sentence Embedding in Learning to Classify Case Sentences
复制标题

DOI:
10.3233/faia210314
复制
发表时间:
2021-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Huihui Xu;Jaromír Šavelka;Kevin D. Ashley
Huihui Xu;Jaromír Šavelka;Kevin D. Ashley
中科院分区:
其他
文献类型:
--
作者:
Huihui Xu;Jaromír Šavelka;Kevin D. Ashley

文献摘要

相似文献

在本文中,我们将句子标注作为一种分类任务。我们采用序列到序列模型来考虑句子位置信息,以确定判例法句子的问题、结论或原因。我们还将法律领域特定句子嵌入与其他通用句子嵌入进行了比较,以衡量在预训练期间捕获的法律领域知识对文本分类的影响。我们将模型部署在摘要和全文决策上。我们发现句子位置信息对全文句子分类特别有用。我们还验证了特定于法律领域的句子嵌入性能更好,并且当包含句子位置信息时,元句嵌入可以进一步提高性能。
In this paper, we treat sentence annotation as a classification task. We employ sequence-to-sequence models to take sentence position information into account in identifying case law sentences as issues, conclusions, or reasons. We also compare the legal domain specific sentence embedding with other general purpose sentence embeddings to gauge the effect of legal domain knowledge, captured during pre-training, on text classification. We deployed the models on both summaries and full-text decisions. We found that the sentence position information is especially useful for full-text sentence classification. We also verified that legal domain specific sentence embeddings perform better, and that meta-sentence embedding can further enhance performance when sentence position information is included.