DT_Team at SemEval-2017 Task 1: Semantic Similarity Using Alignments, Sentence-Level Embeddings and Gaussian Mixture Model Output

DT_Team at SemEval-2017 Task 1: Semantic Similarity Using Alignments, Sentence-Level Embeddings and Gaussian Mixture Model Output
复制标题

DT_Team 在 SemEval-2017 任务 1:使用对齐、句子级嵌入和高斯混合模型输出的语义相似性

DOI:
--
复制
发表时间:
2017
期刊:
International Workshop on Semantic Evaluation
影响因子:
--
通讯作者:
V. Rus
V. Rus
中科院分区:
--
文献类型:
--
作者:
Nabin Maharjan;Rajendra Banjade;D. Gautam;L. Tamang;V. Rus

文献摘要

被引文献

相似文献

我们描述了我们在 SemEval-2017 任务 1、英语语义文本相似性 (STS) 挑战(轨道 5)中提交的系统(DT 团队)。我们开发了三种具有不同功能的不同模型,包括使用单词和块对齐、单词/句子嵌入和高斯混合模型(GMM)计算的相似度分数。我们的系统输出与人类判断之间的相关性高达 0.8536,比基线高出 10% 以上,几乎与性能最佳的系统(相关性为 0.8547)一样好(差异仅为 0.1% 左右)。此外,当使用单独的 STS 基准数据集进行评估时,我们的系统产生了领先的结果。发现基于单词对齐和句子嵌入的特征非常有效。
We describe our system (DT Team) submitted at SemEval-2017 Task 1, Semantic Textual Similarity (STS) challenge for English (Track 5). We developed three different models with various features including similarity scores calculated using word and chunk alignments, word/sentence embeddings, and Gaussian Mixture Model(GMM). The correlation between our system’s output and the human judgments were up to 0.8536, which is more than 10% above baseline, and almost as good as the best performing system which was at 0.8547 correlation (the difference is just about 0.1%). Also, our system produced leading results when evaluated with a separate STS benchmark dataset. The word alignment and sentence embeddings based features were found to be very effective.