CAREER: From One Language to Another
CAREER: From One Language to Another
批准号:
2149404
负责人:
Alexis Palmer
金额:
$55.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-05-15 至 2025-05-31
中文摘要
语言技术已成为我们与信息世界互动的不可或缺的一部分,但复杂的自然语言处理(NLP)工具仅适用于全球约7000种语言中的一小部分。开发自然语言处理工具的现代数据驱动方法一般依赖于有关语言的大量数据的提供,这一障碍对许多语言来说可能是无法克服的,特别是缺乏大量数字资源的语言和发言者人数较少或越来越少的语言。该项目旨在消除为数据较少的语言开发自然语言处理工具的障碍,开发新的方法,将有关语言的语言特性的知识纳入从数据中学习的模型。学习如何建立新语言的NLP工具的更快途径,有可能迅速推动任何语言的语言技术状态。此外,这里开发的工具和知识有可能加快对濒危语言的描述,在仍有说话者可供学习的情况下,帮助确保对世界语言的知情记录。获得语言技术的不平衡出现的部分原因是,目前的自然语言规划模型和算法需要从大量训练数据中学习。这个项目通过采用跨语言迁移学习的方法来解决这种不平衡问题,在跨语言迁移学习中,对一种语言学习的模型进行调整,并利用这些模型对另一种语言进行预测。该项目的创新之一是研究了专家语言知识的融入,以改善模型迁移。两种类型的语言知识将被注入到人工神经网络模型中,用于形态分析和词性标注:a)关于个别语言和语言家族之间的关系的知识;b)关于个别语言和语言家族的特定语言属性的知识。这些模型将进行内在和外在的评估,后者通过研究模型对人类语言分析的有用性,并作为语言文档和描述工作流程的一部分。该奖项反映了NSF的法定使命,并通过使用基金会的智力价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Language technology has become an integral part of how we interact with the world of information, but sophisticated natural language processing (NLP) tools are available only for a handful of the approximately 7000 languages spoken across the world. Modern data-driven methods for developing NLP tools generally rely on the availability of enormous amounts of data for the language in question, an obstacle that may be insurmountable for many languages, especially languages lacking significant digital resources and languages with small or diminishing numbers of speakers. This project aims to remove barriers to developing NLP tools for languages with less data, developing new methods that incorporate knowledge about linguistic properties of languages into models learned from data. Learning how to build faster paths to NLP tools for new languages has the potential to rapidly advance the state of language technology for any language. In addition, the tools and knowledge developed here have the potential to speed up the description of endangered languages, helping to secure an informed record of the world's languages while there are still speakers to learn from.The imbalance in access to language technologies arises in part because current NLP models and algorithms need to learn from large amounts of training data. This project addresses that imbalance by adapting methods from cross-lingual transfer learning, in which models learned on one language are adapted and exploited to make predictions for another language. One innovation of this project is to investigate the incorporation of expert linguistic knowledge for improving model transfer. Two types of linguistic knowledge will be injected into artificial neural network models for morphological analysis and part-of-speech tagging: a) knowledge about relationships between individual languages and language families; and b) knowledge about specific linguistic properties of individual languages and language families. The models will be evaluated both intrinsically and extrinsically, the latter by studying the usefulness of the models for human linguistic analysis and as part of the language documentation and description workflow.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: CCRI: New: Building a Broad Infrastructure for Uniform Meaning Representations
-
批准号:2213805
-
项目类别:Standard Grant
-
资助金额:$100.0万
-
财政年份:2022
-
负责人:Alexis Palmer
-
依托单位:
CAREER: From One Language to Another
-
批准号:1943418
-
项目类别:Continuing Grant
-
资助金额:$55.0万
-
财政年份:2020
-
负责人:Alexis Palmer
-
依托单位:
海外基金