课题基金 / 基金详情

Collaborative Research: CRI: CRD: A Multi-Representational and Multi-Layered Treebank for Hindi/Urdu

Collaborative Research: CRI: CRD: A Multi-Representational and Multi-Layered Treebank for Hindi/Urdu
合作研究:CRI:CRD:印地语/乌尔都语的多表征和多层树库
批准号:
0751213
负责人:
Fei Xia
金额:
$19.6万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-05-01 至 2014-04-30

项目摘要

项目成果

Fei Xia的其他基金

相似基金

相关文献

中文摘要
翻译
树库是自然发生的文本的语料库,这些文本已经用形态和句法(结构)信息进行了注释。在过去的15年里,他们通过为监督机器学习算法提供训练数据,在自然语言处理(NLP)结果方面取得了重大进展。这些算法现在可以自动执行有用的词性标注、解析和语义解释。这个项目正在创建一个新一代的、多代表性的树库。被注释的语言是印地语(40万字)和乌尔都语(20万字)。文本在依赖关系结构(所有节点都用句子的单词标记的树)中进行注释,并使用附加的语义角色标签进行丰富。依赖关系表示也被自动映射到短语结构表示(其中单词位于树的叶子,内部节点用短语标记)。在应用标准的质量控制之后,这两个版本将向公众发布,从而立即提高印地语/乌尔都语NLP的性能。还将发布一个工具,允许研究人员生成短语结构表示的替代格式。这支持将树库视为语言的形态学和语法的更一般、抽象的表示,而不仅仅是特定风格的机器学习实验的数据。对解析和其他NLP任务的研究最近已经认识到重新格式化语法表示以改进机器学习过程的好处;这个树库将使所有对印地语或乌尔都语以及一般语言感兴趣的NLP研究人员更容易迈出这一步。
英文摘要
Treebanks are corpora of naturally occurring text that have been annotated with morphological and syntactic (structural) information. In the last 15 years they have led to significant advances in natural language processing (NLP) results by providing training data for supervised machine learning algorithms. These algorithms can now automatically perform useful part-of-speech tagging, parsing and semantic interpretation. This project is creating a new-generation, multi-representational Treebank. The languages being annotated are Hindi (400K words) and Urdu (200K words). The texts are being annotated in dependency structure (trees in which all nodes are labeled with words of the sentence), enriched with additional semantic role labels. The dependency representation is also being automatically mapped to a phrase-structure representation (in which the words are at the leaves of the tree and internal nodes are labeled with phrase markers). After applying standard quality-control both versions will be released to the public, providing an immediate boost to the performance of Hindi/Urdu NLP. A tool will also be released that will allow a researcher to produce alternative formatting of the phrase structure representation. This supports a view of the treebank as a more general, abstract representation of the morphology and syntax of the language rather than merely as data for a particular style of machine learning experiment. Research into parsing and other NLP tasks has recently recognized the benefits of reformatting syntactic representations in order to improve the machine learning process; this treebank will make that step much easier for all NLP researchers interested in Hindi or Urdu in particular and in language in general.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Workshop on NLP and Linguistics: finding the common ground
  • 批准号:
    1027289
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.7万
  • 财政年份:
    2010
  • 负责人:
    Fei Xia
  • 依托单位:
CAREER: Information Engineering and Synthesis for Resource-poor Languages
  • 批准号:
    0748919
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2008
  • 负责人:
    Fei Xia
  • 依托单位:
CRI:CRD Collaborative Research: General Techniques for Creating Treebanks with Multiple Representations: A Large-Scale Russian
  • 批准号:
    0708719
  • 项目类别:
    Standard Grant
  • 资助金额:
    $2.08万
  • 财政年份:
    2007
  • 负责人:
    Fei Xia
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)