Learning Bill Similarity with Annotated and Augmented Corpora of Bills

Learning Bill Similarity with Annotated and Augmented Corpora of Bills
复制标题

DOI:
10.18653/v1/2021.emnlp-main.787
复制
发表时间:
2021-09
期刊:
--
影响因子:
--
通讯作者:
Jiseon Kim;Elden Griggs;In Song Kim;Alice H. Oh
Jiseon Kim;Elden Griggs;In Song Kim;Alice H. Oh
中科院分区:
其他
文献类型:
--
作者:
Jiseon Kim;Elden Griggs;In Song Kim;Alice H. Oh

文献摘要

相似文献

起草法案是代议制民主的一个重要组成部分。然而,人们往往忽视的是,大多数立法法案都是从其他法案中派生出来的,甚至直接抄袭。尽管法案与法案之间的联系对于理解立法过程具有重要意义,但现有的方法无法解决法案之间的语义相似性,更不用说在法律的文件写作中普遍存在的重新排序或释义了。在本文中,我们克服了这些限制,提出了一个5类分类任务,密切反映了法案生成过程的性质。在此过程中,我们构建了一个包含4,721个分段级别账单到账单关系的人工标记数据集,并将此注释数据集发布给研究社区。为了增强数据集,我们生成具有不同相似度的合成数据,模仿复杂的账单撰写过程。我们使用BERT变体并应用多阶段训练,使用合成和人工标记的数据集依次微调我们的模型。我们发现,当使用人工标记和合成数据进行训练时,预测性能显着提高。最后,我们应用我们的训练模型来推断部分和账单级别的相似性。我们的分析表明,所提出的方法成功地捕捉到的相似性,在不同层次的聚合的法律的文件。
Bill writing is a critical element of representative democracy. However, it is often overlooked that most legislative bills are derived, or even directly copied, from other bills. Despite the significance of bill-to-bill linkages for understanding the legislative process, existing approaches fail to address semantic similarities across bills, let alone reordering or paraphrasing which are prevalent in legal document writing. In this paper, we overcome these limitations by proposing a 5-class classification task that closely reflects the nature of the bill generation process. In doing so, we construct a human-labeled dataset of 4,721 bill-to-bill relationships at the subsection-level and release this annotated dataset to the research community. To augment the dataset, we generate synthetic data with varying degrees of similarity, mimicking the complex bill writing process. We use BERT variants and apply multi-stage training, sequentially fine-tuning our models with synthetic and human-labeled datasets. We find that the predictive performance significantly improves when training with both human-labeled and synthetic data. Finally, we apply our trained model to infer section- and bill-level similarities. Our analysis shows that the proposed methodology successfully captures the similarities across legal documents at various levels of aggregation.