Learning Latent Structures for Cross Action Phrase Relations in Wet Lab Protocols

Learning Latent Structures for Cross Action Phrase Relations in Wet Lab Protocols
复制标题

DOI:
10.18653/v1/2021.acl-long.525
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Chaitanya Kulkarni;Jany Chan;E. Fosler-Lussier;R. Machiraju
Chaitanya Kulkarni;Jany Chan;E. Fosler-Lussier;R. Machiraju
中科院分区:
其他
文献类型:
--
作者:
Chaitanya Kulkarni;Jany Chan;E. Fosler-Lussier;R. Machiraju

文献摘要

相似文献

湿实验室协议(wlp)是在生物学研究中传递可重复性程序的关键。它们由用自然语言编写的指令组成,描述通过特定动作逐步处理材料。wlp中试剂和材料合成的流程描述可以通过材料状态转移图(mstg)来捕获,mstg编码了行动之间的全局时间和因果关系。在这里,我们提出了通过提取多个句子之间的所有动作关系来自动生成给定协议的MSTG的方法。我们还注意到,以前的语料库和方法主要关注动作和实体之间的局部句内关系,而没有解决两个关键问题:(i)解决隐含参数和(ii)建立跨句子的长期依赖关系。我们提出了一个增量学习潜在结构的新模型,它更适合于解决句子间关系和隐含论点。该模型借鉴了一个新的语料库WLP- mstg,该语料库是通过扩展WLP语料库中句子间关系和隐含参数的注释而创建的。我们的模型在我们的语料库中实现了协议的时间和因果关系的F1得分为54.53%,这比以前的模型有了显著的改进——DyGIE++:28.17%;spERT: 27.81%。我们将我们的注释WLP-MSTG语料库提供给研究社区。
Wet laboratory protocols (WLPs) are critical for conveying reproducible procedures in biological research. They are composed of instructions written in natural language describing the step-wise processing of materials by specific actions. This process flow description for reagents and materials synthesis in WLPs can be captured by material state transfer graphs (MSTGs), which encode global temporal and causal relationships between actions. Here, we propose methods to automatically generate a MSTG for a given protocol by extracting all action relationships across multiple sentences. We also note that previous corpora and methods focused primarily on local intra-sentence relationships between actions and entities and did not address two critical issues: (i) resolution of implicit arguments and (ii) establishing long-range dependencies across sentences. We propose a new model that incrementally learns latent structures and is better suited to resolving inter-sentence relations and implicit arguments. This model draws upon a new corpus WLP-MSTG which was created by extending annotations in the WLP corpora for inter-sentence relations and implicit arguments. Our model achieves an F1 score of 54.53% for temporal and causal relations in protocols from our corpus, which is a significant improvement over previous models - DyGIE++:28.17%; spERT:27.81%. We make our annotated WLP-MSTG corpus available to the research community.