When does pretraining help?: assessing self-supervised learning for law and the CaseHOLD dataset of 53,000+ legal holdings

When does pretraining help?: assessing self-supervised learning for law and the CaseHOLD dataset of 53,000+ legal holdings
复制标题

DOI:
10.1145/3462757.3466088
复制
发表时间:
2021-04
期刊:
Proceedings of the Eighteenth International Conference on Artificial Intelligence and Law
影响因子:
--
通讯作者:
Lucia Zheng;Neel Guha;Brandon R. Anderson;Peter Henderson;Daniel E. Ho
Lucia Zheng;Neel Guha;Brandon R. Anderson;Peter Henderson;Daniel E. Ho
中科院分区:
其他
文献类型:
--
作者:
Lucia Zheng;Neel Guha;Brandon R. Anderson;Peter Henderson;Daniel E. Ho

文献摘要

被引文献

相似文献

虽然自监督学习在自然语言处理方面取得了快速进展,但研究人员何时应该进行资源密集型领域特定预训练(域预训练)仍不清楚。令人困惑的是,尽管法律的语言被广泛认为是独特的,但法律却很少有关于领域预训练的实质性收益的记录。我们假设,这些现有的结果源于这样一个事实,即现有的法律的NLP任务太容易,无法满足领域预训练可以帮助的条件。为了解决这个问题,我们首先介绍了CaseHOLD(关于法律的决定的案例持有),这是一个由超过53,000个多项选择题组成的新数据集,用于识别引用案例的相关持有。该数据集为律师提供了一项基本任务,从NLP的角度来看,它既有法律意义,又很困难(F1为0.4,BiLSTM基线)。其次,我们评估CaseHOLD和现有法律的NLP数据集的性能增益。虽然Transformer架构(BERT)在一般语料库上进行预训练,(Google图书和维基百科)提高性能,域预训练使用自定义法律的词汇表(在美国所有法院的1350万个判决的语料库上,这个语料库比BERT的要大),使用CaseHOLD可以获得最大的性能提升(在F1上提高了7.2%,比BERT提高了12%),并在其他两项法律的任务上实现了一致的性能提升。第三,我们表明,域预训练可能是必要的,当任务表现出足够的相似性的预训练语料库:在三个法律的任务的性能增加的水平直接绑定到任务的域特异性。我们的研究结果告知研究人员何时应该进行资源密集型的预训练,并表明基于transformer的架构也会学习暗示不同法律的语言的嵌入。
While self-supervised learning has made rapid advances in natural language processing, it remains unclear when researchers should engage in resource-intensive domain-specific pretraining (domain pretraining). The law, puzzlingly, has yielded few documented instances of substantial gains to domain pretraining in spite of the fact that legal language is widely seen to be unique. We hypothesize that these existing results stem from the fact that existing legal NLP tasks are too easy and fail to meet conditions for when domain pretraining can help. To address this, we first present CaseHOLD (Case Holdings On Legal Decisions), a new dataset comprised of over 53,000+ multiple choice questions to identify the relevant holding of a cited case. This dataset presents a fundamental task to lawyers and is both legally meaningful and difficult from an NLP perspective (F1 of 0.4 with a BiLSTM baseline). Second, we assess performance gains on CaseHOLD and existing legal NLP datasets. While a Transformer architecture (BERT) pretrained on a general corpus (Google Books and Wikipedia) improves performance, domain pretraining (on a corpus of ≈3.5M decisions across all courts in the U.S. that is larger than BERT's) with a custom legal vocabulary exhibits the most substantial performance gains with CaseHOLD (gain of 7.2% on F1, representing a 12% improvement on BERT) and consistent performance gains across two other legal tasks. Third, we show that domain pretraining may be warranted when the task exhibits sufficient similarity to the pretraining corpus: the level of performance increase in three legal tasks was directly tied to the domain specificity of the task. Our findings inform when researchers should engage in resource-intensive pretraining and show that Transformer-based architectures, too, learn embeddings suggestive of distinct legal language.