Intent Classification and Slot Filling for Privacy Policies

Intent Classification and Slot Filling for Privacy Policies
复制标题

DOI:
10.18653/v1/2021.acl-long.340
复制
发表时间:
2021-01
期刊:
--
影响因子:
--
通讯作者:
Wasi Uddin Ahmad;Jianfeng Chi;Tu Le;Thomas B. Norton;Yuan Tian;Kai-Wei Chang
Wasi Uddin Ahmad;Jianfeng Chi;Tu Le;Thomas B. Norton;Yuan Tian;Kai-Wei Chang
中科院分区:
其他
文献类型:
--
作者:
Wasi Uddin Ahmad;Jianfeng Chi;Tu Le;Thomas B. Norton;Yuan Tian;Kai-Wei Chang

文献摘要

相似文献

了解隐私政策对用户来说至关重要,因为它使他们能够了解对他们重要的信息。写在隐私策略文档中的句子解释了隐私实践,组成文本范围传达了关于该实践的进一步具体信息。我们将预测句子中解释的隐私实践称为意图分类,将识别共享特定信息的文本跨度称为槽填充。在这项工作中,我们提出了PolicyIE,这是一个由5250个意图和11788个槽注释组成的英语语料库,涵盖了31个网站和移动应用程序的隐私政策。PolicyIE语料库是一个具有挑战性的现实世界基准,具有有限的标记示例,反映了从领域专家那里收集大规模注释的成本。我们提出了两种替代的神经方法作为基线,(1)意图分类和槽填充作为联合序列标记,(2)将它们建模为序列到序列(Seq2Seq)学习任务。实验结果表明,两种方法在意图分类方面的性能相当,而Seq2Seq方法在槽填充方面的性能明显优于序列标记方法。我们进行了详细的错误分析,以揭示提出的语料库的挑战。
Understanding privacy policies is crucial for users as it empowers them to learn about the information that matters to them. Sentences written in a privacy policy document explain privacy practices, and the constituent text spans convey further specific information about that practice. We refer to predicting the privacy practice explained in a sentence as intent classification and identifying the text spans sharing specific information as slot filling. In this work, we propose PolicyIE, an English corpus consisting of 5,250 intent and 11,788 slot annotations spanning 31 privacy policies of websites and mobile applications. PolicyIE corpus is a challenging real-world benchmark with limited labeled examples reflecting the cost of collecting large-scale annotations from domain experts. We present two alternative neural approaches as baselines, (1) intent classification and slot filling as a joint sequence tagging and (2) modeling them as a sequence-to-sequence (Seq2Seq) learning task. The experiment results show that both approaches perform comparably in intent classification, while the Seq2Seq method outperforms the sequence tagging approach in slot filling by a large margin. We perform a detailed error analysis to reveal the challenges of the proposed corpus.