Person Identification from Text and Speech Genre Samples

Person Identification from Text and Speech Genre Samples
复制标题

DOI:
10.3115/1609067.1609104
复制
发表时间:
2009-03
期刊:
--
影响因子:
--
通讯作者:
Jade Goldstein-Stewart;Ransom K. Winder;R. E. Sabin
Jade Goldstein-Stewart;Ransom K. Winder;R. E. Sabin
中科院分区:
其他
文献类型:
--
作者:
Jade Goldstein-Stewart;Ransom K. Winder;R. E. Sabin

文献摘要

被引文献

相似文献

在本文中,我们描述了一个新的独特的相关语料库的文本和音频样本的人的通信在六个流派上识别一个人进行的实验。文本示例包括散文、电子邮件、博客和聊天。从个人访谈和小组讨论中收集音频样本,然后转录成文本。对于每一种体裁,收集了六个主题的样本。我们表明,我们可以确定的沟通者的准确性为71%的六折交叉验证使用平均22,000字,每个人在六个流派。对于特定类型的人物识别(训练五种类型,测试一种类型),平均准确率达到82%。对于从主题中识别(训练五个主题,测试一个主题),平均准确率达到94%。我们还报告的结果,确定一个人的通信中的一个流派,只使用文本类型,以及音频类型。
In this paper, we describe experiments conducted on identifying a person using a novel unique correlated corpus of text and audio samples of the person's communication in six genres. The text samples include essays, emails, blogs, and chat. Audio samples were collected from individual interviews and group discussions and then transcribed to text. For each genre, samples were collected for six topics. We show that we can identify the communicant with an accuracy of 71% for six fold cross validation using an average of 22,000 words per individual across the six genres. For person identification in a particular genre (train on five genres, test on one), an average accuracy of 82% is achieved. For identification from topics (train on five topics, test on one), an average accuracy of 94% is achieved. We also report results on identifying a person's communication in a genre using text genres only as well as audio genres only.