Extracting Lay Paraphrases of Specialized Expressions from Monolingual Comparable Medical Corpora

Extracting Lay Paraphrases of Specialized Expressions from Monolingual Comparable Medical Corpora
复制标题

从单语可比医学语料库中提取专业表达的通俗释义

DOI:
10.3115/1690339.1690343
复制
发表时间:
2009
影响因子:
4
通讯作者:
Pierre Zweigenbaum
Pierre Zweigenbaum
中科院分区:
心理学2区
文献类型:
--
作者:
Louise Deléger;Pierre Zweigenbaum

文献摘要

被引文献

相似文献

多语种可比语料库已被用于识别单词或术语的翻译,而单语语料库可以帮助识别释义。目前的工作地址之间发现两种不同的话语类型:专业和奠定文本的释义。因此,我们建立了可比语料库的专业和奠定文本,以检测等同奠定和专业的表达。我们确定了两个设备中使用的这样的释义:名词化和新古典化合物。结果表明,这些释义具有很好的精确性,名词化在研究专业语言和非专业语言之间的差异方面确实具有相关性。新古典主义化合物则不那么确定。这项研究还表明,简单的释义获取方法也可以工作的文本具有相当小的相似度,一旦相似的文本片段被检测到。
Whereas multilingual comparable corpora have been used to identify translations of words or terms, monolingual corpora can help identify paraphrases. The present work addresses paraphrases found between two different discourse types: specialized and lay texts. We therefore built comparable corpora of specialized and lay texts in order to detect equivalent lay and specialized expressions. We identified two devices used in such paraphrases: nominalizations and neo-classical compounds. The results showed that the paraphrases had a good precision and that nominalizations were indeed relevant in the context of studying the differences between specialized and lay language. Neo-classical compounds were less conclusive. This study also demonstrates that simple paraphrase acquisition methods can also work on texts with a rather small degree of similarity, once similar text segments are detected.