Creation and Evaluation of Tacit Knowledge Based on Semantic Primes
Creation and Evaluation of Tacit Knowledge Based on Semantic Primes
批准号:
22K12160
负责人:
RZEPKA Rafal
金额:
$2.08万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
2022
资助国家:
日本
项目状态:
未结题
起止时间:
2022-04-01 至 2025-03-31
中文摘要
在第一年里,我一直专注于为进一步的实验开发第一套默会知识。我考虑了有限的语义素数来回答这个问题:在探索环境和理解或学习世界之前,代理应该拥有哪些类型的认知功能?受归入相关语义基元类别的指数及其日语翻译的启发,我创建了一个与知觉相关的23种提示类型的列表。经过一系列的初步实验和注释器测试,我简化、聚合和指定了一些提示,使标注更简短、更容易。为了准备最终的数据集,我使用了以前项目中的日语句子。它的目的是在略有不同的上下文中检测危险级别的变化,由简短的句子对组成,如“孩子吃肥皂”和“孩子吃肥皂形状的糖果”。我创建了一个程序,用于在语义素数的上下文中生成提示,以对代理人、患者和行为进行查询,并聘请了66名注释员选择有关句子的答案。未来实验的最终黄金集由62,687个带注释的句子-提示-选择三元组组成,并在一篇国际会议论文中进行了详细描述,该论文目前正在审查中。实验结果表明,尽管注释者之间的一致性很高,但Bert和Roberta等经典语言模型在基于语义素数的认知感知相关问题识别任务中表现不佳。
英文摘要
During the first year I have concentrated on developing the first set of tacit knowledge for further experiments. I considered the limited number of semantic primes to answer the question "which types of cognitive functionality an agent should posses before exploring the environment and understanding or learning about the world?". Inspired by the exponents grouped into related semantic primitives categories and their translation to Japanese, I created a list of perception-related 23 types of prompts. After series of many preliminary experiments and annotator tests, I simplified, aggregated and specified some prompts to make the annotation shorter and easier.In order to prepare the final dataset I utilized a Japanese sentences from the previous project. Meant for detecting danger level changes in slightly different contexts, it comprises of short sentence pairs as "child eats a soap" and "child eats a soap-shaped candy". I have created a program for generating prompts to make queries about agents, patients and acts in context of semantic primes and hired 66 annotators to choose answers about a sentence.The final golden set for future experiments consisted of 62,687 annotated sentence-prompt-choice triples and has been described in detail in an international conference paper which is currently under review. Experimental results show that although the agreement between annotators was high, classic language models as BERT and RoBERTa performed poorly in a task of recognizing cognitive perception-related questions based on semantic primes.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金