课题基金 / 基金详情

Lexical Chunks and the Nature of Idiolects

Lexical Chunks and the Nature of Idiolects
词汇块和方言的本质
批准号:
2885513
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
我的项目的关键领域是个人风格和作者身份分析。字符n-gram是n个字符的字符串。例如,猫的两个字母是ac、ca和at。字符n-gram似乎是识别作者的一种有用的方法,但对此的解释存在争议。人们并不认为,连续使用一小块字母或单词对一个人来说是与众不同的,但事实确实如此。Grieve et al(2019)等研究结果已经证明了这一点,这些研究结果成功地使用n-gram识别了文本的作者。我将致力于采取基于证据的方法来理解为什么n-gram是有用的。计算机科学家贡献了关于这个主题的大部分研究,但是现在需要更多关于语言理论的专家投入,我希望能做出贡献。我的硕士论文将是一个小规模的研究n-gram跟踪和它占主题的程度。
英文摘要
The key areas of my project are idiolect and authorship analysis. Character n-grams are strings of n characters. For example, the 2-grams in a cat are ac, ca and at. Character n-grams appear to be a useful way of identifying authors, however the explanations for this are disputed. It is not intuitive that the consistent use of small chunks of letters or words could be distinctive to an individual, and yet it can be. This has been exemplified by findings such as that of Grieve et al (2019), that have been successful in identifying the author of a text using n-grams. I am set on my endeavour to undertake an evidence-based approach to understand why n-grams are useful. Computer scientists have contributed the majority of the research about the topic, however now there is a need for greater specialist input regarding linguistic theory, which I hope to contribute to. My master's thesis will be a smaller scale study of n-gram tracing and the extent to which it accounts for topic.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金