CAREER: Information Engineering and Synthesis for Resource-poor Languages
CAREER: Information Engineering and Synthesis for Resource-poor Languages
批准号:
0748919
负责人:
Fei Xia
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-06-15 至 2017-05-31
中文摘要
对于世界上大多数语言来说,语言资源的数量(例如,带注释的语料库和并行数据)非常有限。因此,监督方法和许多非监督方法不能直接应用,使这些语言在很大程度上不受影响和不被注意。自然语言处理(NLP)社区很少关注的另一个关键问题是,迄今为止,很少有研究检查大量语言并将跨语言信息纳入NLP系统。因此,语言被孤立地研究和处理,而不是被视为一个大语系的一部分。这项拟议的研究有两个相互交织的目标。第一个目标是创建一个框架,允许为资源贫乏的语言快速开发资源。通过将语法信息从资源丰富的语言投射到资源贫乏的语言,引导NLP工具使用初始种子来实现这一目标。第二个目标是利用自动创建的资源,对大量语言进行跨语言研究,发现语言知识。这些知识不仅会加深我们对语言的理解,而且还提供了额外的信息,可以纳入引导模块,以产生更好的NLP工具。该研究探索了两个关键的想法:第一个想法是利用资源丰富的语言,通过使用它们来创建引导NLP工具的种子。第二个想法是识别语言之间的关系,并使用这些信息来帮助机器学习。两种观点都指向同一个方向;也就是说,语言是相互关联的,应该这样对待。
英文摘要
For the majority of the world's languages, the amount of linguistic resources (e.g., annotated corpora and parallel data) is very limited. Consequently, supervised methods and many unsupervised methods cannot be applied directly, leaving these languages largely untouched and unnoticed. Another crucial issue, which has received little attention from the natural language processing (NLP) community, is that to date there have been very few studies that examine a large number of languages and incorporate cross-lingual information into NLP systems. As a result, languages are researched and processed in isolation rather than being looked at as part of a big language family.This proposed research has two intertwined goals. The first goal is to create a framework that allows the rapid development of resources for resource-poor languages. This goal will be accomplished by bootstrapping NLP tools with initial seeds created by projecting syntactic information from resource-rich languages to resource-poor ones. The second goal is to use the automatically created resources to perform cross-lingual study on a large number of languages to discover linguistic knowledge. The knowledge will not only deepen our understanding on languages, but also provide additional information that can be incorporated into the bootstrapping module to produce better NLP tools. The research explores two key ideas: The first idea is to take advantage of resource-rich languages by using them to create seeds for bootstrapping NLP tools. The second idea is to identify the relation between languages and use this information to help machine learning. Both ideas point to the same direction; that is, languages are related to one another and should be treated as such.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Workshop on NLP and Linguistics: finding the common ground
-
批准号:1027289
-
项目类别:Standard Grant
-
资助金额:$1.7万
-
财政年份:2010
-
负责人:Fei Xia
-
依托单位:
Collaborative Research: CRI: CRD: A Multi-Representational and Multi-Layered Treebank for Hindi/Urdu
-
批准号:0751213
-
项目类别:Continuing Grant
-
资助金额:$19.6万
-
财政年份:2008
-
负责人:Fei Xia
-
依托单位:
CRI:CRD Collaborative Research: General Techniques for Creating Treebanks with Multiple Representations: A Large-Scale Russian
-
批准号:0708719
-
项目类别:Standard Grant
-
资助金额:$2.08万
-
财政年份:2007
-
负责人:Fei Xia
-
依托单位:
国内基金
海外基金
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
-
批准号:--
-
项目类别:外国青年学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:江洋子
-
依托单位:
Exploring the Intrinsic Mechanisms of CEO Turnover and Market Reaction: An Explanation Based on Information Asymmetry
-
批准号:W2433169
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:HAOFEI ZHANG
-
依托单位:
SCIENCE CHINA Information Sciences
-
批准号:61224002
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:宋扉
-
依托单位: