CAREER: Information Engineering and Synthesis for Resource-poor Languages
CAREER: Information Engineering and Synthesis for Resource-poor Languages
批准号:
0748919
负责人:
Fei Xia
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-06-15 至 2017-05-31
中文摘要
对于世界上大多数语言来说,语言资源的数量(例如,注释的语料库和并行数据)非常有限。因此,有监督的方法和许多无监督的方法不能直接应用,使这些语言在很大程度上不受影响和忽视。自然语言处理(NLP)社区很少关注的另一个关键问题是,到目前为止,很少有研究考察大量语言并将跨语言信息纳入NLP系统。因此,语言的研究和处理是孤立的,而不是被视为一个大的语言家族的一部分。第一个目标是创建一个框架,允许快速开发资源贫乏语言的资源。这一目标将通过引导NLP工具来实现,初始种子是通过将语法信息从资源丰富的语言投射到资源贫乏的语言而创建的。第二个目标是使用自动创建的资源对大量语言进行跨语言学习,以发现语言知识。这些知识不仅可以加深我们对语言的理解,还可以提供额外的信息,这些信息可以纳入引导模块,以产生更好的NLP工具。该研究探讨了两个关键思想:第一个思想是利用资源丰富的语言,使用它们来创建自举NLP工具的种子。第二个想法是识别语言之间的关系,并使用这些信息来帮助机器学习。这两种观点都指向同一个方向,即语言是相互关联的,应该这样对待。
英文摘要
For the majority of the world's languages, the amount of linguistic resources (e.g., annotated corpora and parallel data) is very limited. Consequently, supervised methods and many unsupervised methods cannot be applied directly, leaving these languages largely untouched and unnoticed. Another crucial issue, which has received little attention from the natural language processing (NLP) community, is that to date there have been very few studies that examine a large number of languages and incorporate cross-lingual information into NLP systems. As a result, languages are researched and processed in isolation rather than being looked at as part of a big language family.This proposed research has two intertwined goals. The first goal is to create a framework that allows the rapid development of resources for resource-poor languages. This goal will be accomplished by bootstrapping NLP tools with initial seeds created by projecting syntactic information from resource-rich languages to resource-poor ones. The second goal is to use the automatically created resources to perform cross-lingual study on a large number of languages to discover linguistic knowledge. The knowledge will not only deepen our understanding on languages, but also provide additional information that can be incorporated into the bootstrapping module to produce better NLP tools. The research explores two key ideas: The first idea is to take advantage of resource-rich languages by using them to create seeds for bootstrapping NLP tools. The second idea is to identify the relation between languages and use this information to help machine learning. Both ideas point to the same direction; that is, languages are related to one another and should be treated as such.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Workshop on NLP and Linguistics: finding the common ground
-
批准号:1027289
-
项目类别:Standard Grant
-
资助金额:$1.7万
-
财政年份:2010
-
负责人:Fei Xia
-
依托单位:
Collaborative Research: CRI: CRD: A Multi-Representational and Multi-Layered Treebank for Hindi/Urdu
-
批准号:0751213
-
项目类别:Continuing Grant
-
资助金额:$19.6万
-
财政年份:2008
-
负责人:Fei Xia
-
依托单位:
CRI:CRD Collaborative Research: General Techniques for Creating Treebanks with Multiple Representations: A Large-Scale Russian
-
批准号:0708719
-
项目类别:Standard Grant
-
资助金额:$2.08万
-
财政年份:2007
-
负责人:Fei Xia
-
依托单位:
国内基金
海外基金
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
-
批准号:--
-
项目类别:外国青年学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:江洋子
-
依托单位:
Exploring the Intrinsic Mechanisms of CEO Turnover and Market Reaction: An Explanation Based on Information Asymmetry
-
批准号:W2433169
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:HAOFEI ZHANG
-
依托单位:
SCIENCE CHINA Information Sciences
-
批准号:61224002
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:宋扉
-
依托单位: