RI: Small: Modeling Lexical Borrowing to Bridge the "Linguistic Divide" in Natural Language Processing
RI: Small: Modeling Lexical Borrowing to Bridge the "Linguistic Divide" in Natural Language Processing
批准号:
1526745
负责人:
Alan Black
金额:
$45.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2018-08-31
中文摘要
丰富的智能语言感知技术生态系统(例如,个人助理、内容推荐、垃圾邮件检测等)英语和其他高资源语言的用户可以访问的资源取决于是否存在特定于语言的数据资源。开发支持这些技术的资源通常需要大量投资--无论是在金钱上还是在训练有素的母语人士方面--这意味着如果没有新的战略,世界上7000多种语言中的大多数可能仍然是资源匮乏,其母语者得不到充分的服务。该项目通过确定高资源和低资源语言之间的跨语言对应并相应地投射资源(例如,翻译、词汇本体和句法注释),更经济地解决了以低资源语言引导语言技术所需的语言资源的问题。为了识别这些对应关系,这项工作发展了语言借用的计算模型,即由于语言接触和双语能力,来自捐赠者语言的单词被接受者顺应的过程。除了能够将资源从资源丰富的语言转移到资源匮乏的语言之外,能够识别借词还能够基于语料库对已被确定为与借词相关的社会因素(国家之间的权力差异、公众舆论和地理位置、性别和种族/民族等个人属性)进行研究。因此,通过观察语言的变化,这项工作使社会关系的变化能够被量化。借词的过程并不是一成不变的,对这一过程进行建模是识别借词实例的核心挑战。幸运的是,自适应过程通常是规则的,并服从计算建模,这项工作使用了加权有限状态换能器,其特征来自最优理论(OT)。与传统的语言学幼稚的统计模型相比,OT派生的特征不仅提供了更高的统计效率,而且它们还提供了一种基于语料库的新的基于语料库的验证音系理论的一些核心主张。借用模型确定了数十个具有类型代表性的语言对的词汇对应关系(主要文本数据从维基百科、Twitter、博客和在线新闻等开放资源中获得),从而能够预测资源和开发核心自然语言处理技术。最后,借词模型能够在文本中识别随着时间的推移而产生的借词实例,从而实现基于语料库的社会语言学研究。
英文摘要
The rich ecosystem of intelligent, language-aware technologies (e.g., personal assistants, content recommendation, spam detection, etc.) that users of English and other high-resource languages have access to depends on the existence of language-specific data resources. Developing the resources that enable these technologies has usually required a substantial investment -- both monetarily and in terms of trained native speakers -- meaning that without new strategies, most of the 7,000+ languages in the world would likely remain resource-poor and their speakers underserved. This project addresses the problem of bootstrapping linguistic resources required for language technologies in low-resource languages more economically by identifying cross-linguistic correspondences between high- and low-resource languages and projecting resources (e.g., translations, lexical ontologies, and syntactic annotations) accordingly. To identify these correspondences, this work develops computational models of linguistic borrowing, which is the process by which words from a donor language are adapted by speakers of a recipient language as a result of language contact and bilingualism. In addition to enabling the transfer of resources from high- to low-resource languages, being able to identify borrowing enables corpus-based studies of the social factors (power differences between countries, public opinion, and personal attributes such as geographic location, gender, and race/ethnicity) that have been identified as correlates with which words are borrowed. Thus, by observing language change, this work enables changes in social relations to be quantified.Words are not left unchanged by the process of borrowing, and modeling this process is the central challenge to identifying instances of borrowing. Fortunately, the adaptation processes are generally regular and amenable to computational modeling, and this work uses weighted finite-state transducers parameterized with features derived from Optimality Theory (OT). OT-derived features not only provide increased statistical efficiency relative to conventional linguistically naive statistical models but they also provide a new kind of corpus-based verification of some of the central claims of phonological theory. The borrowing model identifies lexical correspondences across dozens of typologically representative language pairs (primary text data is obtained from open resources such as Wikipedia, Twitter, blogs, and online news), enabling projection of resources and development of core natural language processing technologies. Finally, the borrowing model enables instances of borrowed words to be identified in text as it is generated over time, enabling corpus-based sociolinguistic studies.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
TRANSFORM: flexible voice synthesis through articulatory voice transformation
-
批准号:0414675
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Alan Black
-
依托单位:
ITR: Evaluation and Personalization of Synthetic Voices
-
批准号:0219687
-
项目类别:Continuing Grant
-
资助金额:$39.43万
-
财政年份:2002
-
负责人:Alan Black
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: