课题基金 / 基金详情

RI: Small: Demographic-Aware Lexical Semantics

RI: Small: Demographic-Aware Lexical Semantics
RI:小:人口感知词汇语义
批准号:
1815291
负责人:
Rada Mihalcea
金额:
$45.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-09-01 至 2021-08-31

项目摘要

项目成果

Rada Mihalcea的其他基金

相似基金

相关文献

中文摘要
翻译
自然语言处理中的一个核心挑战是开发确定单词的含义如何相互关联的方法。 这个任务被称为“词汇语义学”,因为“词汇”意味着“词”,“语义学”意味着“意义”。 传统的词典并没有解决词汇语义的问题,因为定义往往是循环的或不完整的,特别是对于最常见的单词。 相反,词汇语义模型是通过处理大量文本来计算的,使用的原则是,经常出现在相同上下文中的成对单词在某些维度上必须具有相似的含义。 例如,“man”和“boy”这两个词可被推断为具有与“human”和“gender”维度沿着的含义。 然而,当前模型的一个局限性是,它们假设单词的含义对于一种语言的所有使用者都是相同的。 这显然是错误的:例如,我们知道,讲英语的人使用不同的单词取决于他们的年龄、性别、工作领域和地理位置等因素,也就是说,基于他们的人口统计数据。 本项目将通过开发人口统计感知词汇语义的方法来克服这一限制,其中以人为本的信息补充了基于语言的信息。 这项工作将有助于改善人与计算机之间的自然语言通信系统,如Siri或Alexa,以及改善不同语言之间的自动翻译系统。近年来,使用基于语料库的方法(如分布式向量空间模型和词嵌入)进行词汇语义研究取得了重大进展。与此同时,Web 2.0的发展带来了大量的文本,其中大多数都含有丰富的显性或隐性人口统计信息,如作者的年龄,性别,行业或位置。该项目的目标是在这两种趋势的交汇处采取下一个自然步骤,并开发人口统计感知词汇语义的方法,其中以人为本的信息补充了基于语言的信息,以增强语言表示,明确说明语言背后的人口统计和特征。 该项目主要有以下三个研究目标。首先,它开发了新的人口统计感知的单词表示模型,不仅考虑上下文知识,而且以人为本的信息。探索的方法包括分布式向量空间模型,可以组成,以创建人口统计学感知的向量空间表示为各种人口概况,联合词嵌入,结合联合收割机通用的基于上下文的嵌入与专门的嵌入,反映了特定的人口维度。第二,建立在广泛的行为研究的目标是识别系统的异质性跨群体的先前工作,实验室研究的设计,以验证从计算模型的结果。 第三,这些新的以人为本的词表示在自然语言处理的三个核心任务的应用进行了探索,从简单到复杂,即:词联想,文本相似性,和多样化的news.This奖项反映了NSF的法定使命,并已被认为是值得通过使用基金会的智力价值和更广泛的影响力审查标准进行评估的支持。
英文摘要
A central challenge in natural language processing is to develop methods for determining how meanings of words relate to one another. This task is called "lexical semantics", because "lexical" means "word" and "semantics" means "meaning". Traditional dictionaries do not solve the problem of lexical semantics, because definitions are often circular or incomplete, especially for the most common words. Instead, models of lexical semantics are computed by processing large bodies of text, using the principle that pairs of words that often appear in the same contexts must have meanings that are similar along some dimensions. For example, the words "man" and "boy" would be inferred to have similar meanings along the dimensions of "human" and "gender". However, a limitation of current models is that they assume that the meaning of words is the same for all speakers of a language. This is plainly false: we know, for example, that English speakers use words differently depending, among other factors, their age, gender, field of work, and geographic location; that is, on the basis of their demographics. This project will overcome this limitation by developing methods for demographic-aware lexical semantics, where people-centric information complements language-based information. This work will help improve systems for natural language communication between people and computers, such as Siri or Alexa, as well as improve systems for automatically translating between different languages.Recent years have witnessed significant progress in research in lexical semantics using corpus-based approaches such as distributional vector-space models and word embeddings. At the same time, the growth of Web 2.0 has led to tremendous volumes of texts, most of which are rich in explicit or implicit demographic information, such as the age, gender, industry, or location of the writer. The goal of this project is to take the next natural step at the confluence of these two trends, and develop methods for demographic-aware lexical semantics, where people-centric information complements language-based information for enhanced linguistic representations that explicitly account for the demographics and traits of the people behind the language. The project targets the following three main research objectives. First, it develops novel demographic-aware word representations models that account not only for contextual knowledge but also for people-centric information. Methods that are explored include distributional vector-space models that can be composed to create demographic-aware vector-space representations for various demographic profiles, and joint word embeddings that combine generic context-based embeddings with specialized embeddings that reflect the specifics of given demographic dimensions. Second, building upon extensive previous work in behavioral studies targeting the identification of systematic heterogeneity across groups, lab studies are devised to validate the findings from the computational models. Third, the application of these novel people-centric word representations to three core tasks in natural language processing are explored, ranging from simple to complex, namely: word associations, text similarity, and diversified news.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
Measuring Personal Values in Cross-Cultural User-Generated Content
衡量跨文化用户生成内容中的个人价值观
DOI: --
发表时间: 2019
期刊: International Conference on Social Informatics
影响因子: --
作者: [Yiting Shen, Steven R.]
通讯作者: Yiting Shen, Steven R.
DOI: 10.18653/v1/2021.socialnlp-1.13
发表时间: 2021-05
期刊: ArXiv
影响因子: --
作者: [MeiXing Dong;Xueming Xu;Yiwei Zhang;Ian Stewart;Rada Mihalcea]
通讯作者: MeiXing Dong;Xueming Xu;Yiwei Zhang;Ian Stewart;Rada Mihalcea
DOI: 10.18653/v1/2021.emnlp-main.476
发表时间: 2020-04
期刊:
影响因子: --
作者: [Laura Burdick;Jonathan K. Kummerfeld;Rada Mihalcea]
通讯作者: Laura Burdick;Jonathan K. Kummerfeld;Rada Mihalcea
DOI: 10.1109/mis.2019.2916965
发表时间: 2019-07
期刊: IEEE Intelligent Systems
影响因子: 6.4
作者: [C. Welch;Verónica Pérez-Rosas;Jonathan K. Kummerfeld;Rada Mihalcea;E. Cambria]
通讯作者: C. Welch;Verónica Pérez-Rosas;Jonathan K. Kummerfeld;Rada Mihalcea;E. Cambria
共 13 条
    SCH: Natural Language Processing for Enhanced Behavioral Counseling
    CAREER: Semantic Interpretation with Monolingual and Cross-lingual Evidence
    INSPIRE Track 1: Language-Based Computational Methods for Analyzing Worldviews
    RI: Small: Collaborative Research: Word Sense and Multilingual Subjectivity Analysis
    • 批准号:
      0917170
    • 项目类别:
      Standard Grant
    • 资助金额:
      $22.48万
    • 财政年份:
      2009
    • 负责人:
      Rada Mihalcea
    • 依托单位:
    国内基金
    海外基金
    昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
    • 依托单位:
    tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      10.0万元
    • 批准年份:
      2022
    • 负责人:
      张祥忠
    • 依托单位:
    Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
    Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
    • 批准号:
      31972324
    • 项目类别:
      面上项目
    • 资助金额:
      58.0万元
    • 批准年份:
      2019
    • 负责人:
      高学文
    • 依托单位: