课题基金 / 基金详情

BIGDATA: Small: DA: Big Multilinguality for Data-Driven Lexical Semantics

BIGDATA: Small: DA: Big Multilinguality for Data-Driven Lexical Semantics
BIGDATA:小:DA:数据驱动词汇语义的大多语言性
批准号:
1251131
负责人:
Noah Smith
金额:
$25.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-08-01 至 2015-09-30

项目摘要

项目成果

Noah Smith的其他基金

相似基金

相关文献

中文摘要
翻译
自然语言处理中的一个关键挑战是定义单词的计算表示。数据驱动的分布方法使用语料库根据单词出现的上下文来归纳单词的向量空间表示。该项目超越了传统的方法(例如,潜在语义分析;Deerwester等人,1990),后者使用语料库中往往出现在单词附近的单词来定义上下文,方法是扩展用于构建语义向量的上下文类型。首先,该项目除了传统的单语语料库外,还纳入了翻译语境,即多语种平行语料库中容易获得的单词。这允许跨语言共享证据,最重要的是,从拥有大型语料库的资源丰富的语言到资源更贫乏的语言。其次,该项目结合了通过作者、时间、地理和社交连接元数据捕获的可从社交网络平台推断的社交上下文。综上所述,这些额外的功能为单词的上下文提供了更广泛的定义,并导致了对人类语言建模的分布式方法的更统一方法,朝着独立于语言的语义方向发展。该项目侧重于代表几个主要语系的十种不同类型的语言(英语、阿拉伯语、汉语、西班牙语、俄语、德语、葡萄牙语、斯瓦希里语、马达加斯加语和波斯语)。一个关键的重点是扩大算法,以推断网络规模的语料库的分布表示,并处理代表扩展的上下文概念的更大的上下文向量。该方法还利用嘈杂的句法处理,使得在定义上下文时能够考虑句法信息,而不仅仅是关于相邻单词的信息。除了通过包括更丰富的上下文信息来提高学习的词汇语义表示的质量之外,该项目还创建了链接跨语言的单词类型的词汇语义表示。它们可以直接用于文本处理应用,例如文本分类、机器翻译、信息提取和文本的语义分析,并且它们将能够以资源较少的语言构建健壮的词汇语义资源,这些语言受益于它们所配对的语言中丰富的资源。制作的多语种矢量表示将发布给研究社区,并将用于本科生的课堂项目。该项目为两名研究生在动态的研究环境中提供综合的教育和研究体验。项目网站(http://www.ark.cs.cmu.edu/BigMultilinguality)将用于传播成果。
英文摘要
A key challenge in natural language processing is defining the computational representation of words. Data-driven distributional approaches use corpora to induce vector-space representations for words, based on the contexts they occur in. This project goes beyond traditional approaches (e.g., latent semantic analysis; Deerwester et al., 1990), which use words that tend to occur near a word in corpora to define the context, by extending the types of contexts used in constructing semantic vectors. First, this project incorporates translation contexts, i.e., words readily available in multilingual parallel corpora, alongside traditional monolingual corpora. This allows evidence-sharing across languages, most importantly from resource-rich languages with large corpora to more resource-poor languages. Second, this project incorporates social context inferable from social network platforms, captured through author, time, geographic, and social connection metadata. Taken together, these additional features give a broader definition of a word's context and lead to a more unified approach to the distributional approach to modeling human language, moving in the direction of a language-independent semantics. The project focuses on ten typologically diverse languages representing several major language families (English, Arabic, Chinese, Spanish, Russian, German, Portuguese, Swahili, Malagasy, and Farsi). A key emphasis is scaling up algorithms for inferring distributional representations to web-scale corpora and dealing with much larger contextual vectors representing the expanded notion of context. The approach also leverages noisy syntactic processing to enable syntactic information, rather than just information about neighboring words, to be considered when defining context.In addition to improving the quality of the learned lexico-semantic representations by including richer contextual information, this project creates lexical semantic representations that link word types across languages. These have direct use in text processing applications such as text categorization, machine translation, information extraction, and semantic analysis of text, and they will enable the construction of robust lexical semantic resources in lower-resource languages that benefit from the richness of resources in languages they are paired with. The multilingual vector representations produced will be released to the research community and will be used in undergraduate class projects. The project provides integrated educational and research experience for two graduate students in a dynamic research environment. The project website (http://www.ark.cs.cmu.edu/BigMultilinguality) will be used for dissemination of results.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
NSF-BSF: RI: Small: Efficient Transformers via Formal and Empirical Analysis
  • 批准号:
    2113530
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.98万
  • 财政年份:
    2021
  • 负责人:
    Noah Smith
  • 依托单位:
RI/SES: Conference Proposal: Doctoral Consortium on Text as Data
  • 批准号:
    1830158
  • 项目类别:
    Standard Grant
  • 资助金额:
    $2.5万
  • 财政年份:
    2018
  • 负责人:
    Noah Smith
  • 依托单位:
NSF-BSF: RI: Small: Collaborative Research: Modeling Crosslinguistic Influences Between Language Varieties
  • 批准号:
    1813153
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $16.75万
  • 财政年份:
    2018
  • 负责人:
    Noah Smith
  • 依托单位:
RI: Medium: Broad-Coverage Semantic Parsing: Linguistic Representation Learning from Crowd-Scale Data
  • 批准号:
    1562364
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $100.6万
  • 财政年份:
    2016
  • 负责人:
    Noah Smith
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: