Collaborative Research: CRI: CRD: A Multi-Representational and Multi-Layered Treebank for Hindi/Urdu
Collaborative Research: CRI: CRD: A Multi-Representational and Multi-Layered Treebank for Hindi/Urdu
批准号:
0751213
负责人:
Fei Xia
金额:
$19.6万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-05-01 至 2014-04-30
中文摘要
树库是自然出现的文本的语料库,已经用形态和句法(结构)信息进行了注释。在过去的15年里,它们通过为有监督的机器学习算法提供训练数据,在自然语言处理(NLP)结果方面取得了显着的进步。这些算法现在可以自动执行有用的词性标注、句法分析和语义解释。这个项目正在创建一个新一代、多代表性的树库。被注释的语言是印地语(400K字)和乌尔都语(200K字)。文本在依存结构(其中所有节点都用句子的单词标记的树)中进行标注,并使用额外的语义角色标签进行丰富。依存关系表示也被自动映射到短语结构表示(其中单词位于树的叶子处,并且内部节点用短语标记标记)。在应用标准质量控制后,两个版本都将向公众发布,立即提高印地语/乌尔都语NLP的性能。还将发布一个工具,使研究人员能够产生短语结构表示的替代格式。这支持了一种观点,即树库是语言的词法和句法的更一般、更抽象的表示,而不仅仅是特定风格的机器学习实验的数据。对语法分析和其他NLP任务的研究最近认识到了重新格式化句法表示以改进机器学习过程的好处;这个树库将使所有对印地语或乌尔都语特别感兴趣的NLP研究人员以及一般语言研究人员更容易完成这一步骤。
英文摘要
Treebanks are corpora of naturally occurring text that have been annotated with morphological and syntactic (structural) information. In the last 15 years they have led to significant advances in natural language processing (NLP) results by providing training data for supervised machine learning algorithms. These algorithms can now automatically perform useful part-of-speech tagging, parsing and semantic interpretation. This project is creating a new-generation, multi-representational Treebank. The languages being annotated are Hindi (400K words) and Urdu (200K words). The texts are being annotated in dependency structure (trees in which all nodes are labeled with words of the sentence), enriched with additional semantic role labels. The dependency representation is also being automatically mapped to a phrase-structure representation (in which the words are at the leaves of the tree and internal nodes are labeled with phrase markers). After applying standard quality-control both versions will be released to the public, providing an immediate boost to the performance of Hindi/Urdu NLP. A tool will also be released that will allow a researcher to produce alternative formatting of the phrase structure representation. This supports a view of the treebank as a more general, abstract representation of the morphology and syntax of the language rather than merely as data for a particular style of machine learning experiment. Research into parsing and other NLP tasks has recently recognized the benefits of reformatting syntactic representations in order to improve the machine learning process; this treebank will make that step much easier for all NLP researchers interested in Hindi or Urdu in particular and in language in general.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Workshop on NLP and Linguistics: finding the common ground
-
批准号:1027289
-
项目类别:Standard Grant
-
资助金额:$1.7万
-
财政年份:2010
-
负责人:Fei Xia
-
依托单位:
CAREER: Information Engineering and Synthesis for Resource-poor Languages
-
批准号:0748919
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2008
-
负责人:Fei Xia
-
依托单位:
CRI:CRD Collaborative Research: General Techniques for Creating Treebanks with Multiple Representations: A Large-Scale Russian
-
批准号:0708719
-
项目类别:Standard Grant
-
资助金额:$2.08万
-
财政年份:2007
-
负责人:Fei Xia
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: