课题基金 / 基金详情

Author Profiling using stylometry

Author Profiling using stylometry
使用文体测量法进行作者分析
批准号:
2881667
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该项目旨在进一步通过语言进行分析的科学,以便一旦检索到书面样本,就可以快速有效地进行分析。因此,我们将专注于分析作者的任务:目标是从他们的写作中提取风格信号,特定于语言社区的模式,看看当该社区的成员在该语言社区之外的环境中写作时,这些信号是否可以检测到。这背后的理论是,我们都参与了不同的语言社区,我们的假设是,我们所处的每个语言社区都会在我们的写作和说话风格中留下印记,其中一些印记可以通过文体测量学来检测,文体测量学是对写作风格的定量研究。为了实现这一点,我们将首先看看已经完成的文体分析任务,以及他们的方法的成功,以便对剖析者已经可以使用的工具以及如何使用这些工具提供细致入微的总结。这样做还可以让我们了解分析社区的需求,以便创建一个优先级列表,并将其转化为我们进行的实验。我们着手的每个分析任务都很可能需要一个新的语料库,并有自己的策展需求,因为我们必须确保最大限度地减少混淆变量。如果得到适当的维护和更新,我们创建的语料库也可以为其他分析人员提供服务,(通过交叉验证和我们的实验)为特定的分析任务工作。为了减轻每个分析任务的风险,我们将逐步收集语料库,以便定期检查成功和准确性,使我们能够始终如一地报告项目的进展情况,并决定哪些任务是可行的。
英文摘要
This project aims to further the science of profiling through language so that it can be done quickly and effectively as soon as a written sample is retrieved. We will therefore focus on the task of profiling an author: The goal is to extract stylistic signals, patterns specific to a linguistic community from their writing, to see if these signals are detectable when members of that community write in a context outside of that linguistic community. The theory behind this is that we all take part in different linguistic communities, and our assumption is that each linguistic community we are a part of leaves a mark in our writing and speaking style, some of which may be detectable using stylometry, the quantitative study of writing style.In order to achieve this, we will first take a look at the stylometric profiling tasks that have already been done, and the success in their methodology, in order to provide a nuanced summary of the tools profilers can already have at their disposal and how to use them. Doing this will also allow us to understand the needs of the profiling community, in order to create a list of priorities that will translate into experiments we carry out.Each profiling task we embark on will most likely require a new corpus with its own curation needs, as we must make sure to minimize confounding variables. If properly maintained and updated, the corpora we create can also serve for other profilers to carry out their work with a corpus that is known (through cross-validation and our experimentation) to work for a particular profiling task.To mitigate the risks for each profiling task, we will gather the corpora incrementally, so as to have regular checks for success and accuracy that will allow us to consistently make reports of the project's progress and decide which tasks are feasible.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金