Transformer-Based Approaches for Legal Text Processing

Transformer-Based Approaches for Legal Text Processing
复制标题

DOI:
10.1007/s12626-022-00102-2
复制
发表时间:
2022-01
期刊:
The Review of Socionetwork Strategies
影响因子:
--
通讯作者:
Nguyen Ha Thanh;Phuong Minh Nguyen;Thi-Hai-Yen Vuong;Minh Q. Bui;Minh-Chau Nguyen;Binh Dang;Vu Tran;Le-Minh Nguyen;Kenji Satoh
Nguyen Ha Thanh;Phuong Minh Nguyen;Thi-Hai-Yen Vuong;Minh Q. Bui;Minh-Chau Nguyen;Binh Dang;Vu Tran;Le-Minh Nguyen;Kenji Satoh
中科院分区:
其他
文献类型:
--
作者:
Nguyen Ha Thanh;Phuong Minh Nguyen;Thi-Hai-Yen Vuong;Minh Q. Bui;Minh-Chau Nguyen;Binh Dang;Vu Tran;Le-Minh Nguyen;Kenji Satoh

文献摘要

被引文献

相似文献

本文介绍了我们针对COLIEE 2021自动法律文本处理竞赛中的不同问题所采用的基于Transformer模型的方法。由于法律文件的特点以及数据量的限制,法律文件的自动化处理是一项具有挑战性的任务。通过详细的实验,我们发现,通过适当的方法,基于Transformer的预训练语言模型可以很好地处理自动法律文本处理问题。我们详细描述了每个任务的处理步骤,如问题形成、数据处理和增强、预训练、精调。此外,我们还向社区介绍了两个利用法律领域并行翻译的预训练模型:NFSP和NMSP。其中,NFSP在大赛任务5中取得了最先进的成绩。虽然本文关注的是技术报告,但其方法的新颖性也可以为使用基于Transformer的模型的自动法律文档处理提供有用的参考。
In this paper, we introduce our approaches using Transformer-based models for different problems of the COLIEE 2021 automatic legal text processing competition. Automated processing of legal documents is a challenging task because of the characteristics of legal documents as well as the limitation of the amount of data. With our detailed experiments, we found that Transformer-based pretrained language models can perform well with automated legal text-processing problems with appropriate approaches. We describe in detail the processing steps for each task such as problem formulation, data processing and augmentation, pretraining, finetuning. In addition, we introduce to the community two pretrained models that take advantage of parallel translations in legal domain, NFSP and NMSP. In which, NFSP achieves the state-of-the-art result in Task 5 of the competition. Although the paper focuses on technical reporting, the novelty of its approaches can also be an useful reference in automated legal document processing using Transformer-based models.