课题基金 / 基金详情

Interactive distributed corpus exploration and annotation infrastructure for large corpora and knowledge-bases

Interactive distributed corpus exploration and annotation infrastructure for large corpora and knowledge-bases
适用于大型语料库和知识库的交互式分布式语料库探索和注释基础设施
批准号:
315979217
负责人:
Dr.-Ing. Richard Eckart de Castilho
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research data and software (Scientific Library Services and Information Systems)
财政年份:
2016
资助国家:
德国
项目状态:
已结题
起止时间:
2015-12-31 至 2021-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
这个项目的目标是为语料库标注提供一个研究基础设施,通过灵活地构建子语料库来扩展到大型文本文档集合。该基础设施满足了计算语言学家和语料库语言学家对通用工具的需求,以便在文档内部和跨文档执行选择性语义注释任务。这样的基础设施很重要,因为它可以有针对性地利用大量的数字文本进行语言分析。专家用户在探索大型文档集合、建立注释方案以及从大型背景语料库中提取特定任务的子语料库方面应该得到基础设施的支持。语料库的注释应该灵活地分配给不同资格水平和背景的远程工作注释团队。他们的工作应该通过基于机器学习技术的优先级和注释建议来支持,以有效地创建具有高质量注释的大型语料库,用于训练和评估各自的算法。因此,基础设施应该允许多个研究人员和并行工作的注释团队从多个角度对相同的语料库进行注释。用户应该根据需要导入自定义语料库。还需要进一步的功能来维护和扩展语义注释任务期间使用的知识库,并连接到外部标准知识库。
英文摘要
The goal of this project is a research infrastructure for corpus annotation that scales to large text document collections by flexibly building subcorpora. The infrastructure addresses the needs of computational linguists and corpus linguists for a generic tool to perform selective semantic annotation tasks within and across documents. Such an infrastructure is important because it enables the targeted exploitation of the huge amounts of digitally available text for linguistic analysis. The expert user should be supported by the infrastructure in exploring the large document collections, in setting up an annotation scheme, and in extracting task-specific subcorpora from a large background corpus. The annotation of the corpora should be flexibly distributable to remotely working annotation teams of different qualification levels and backgrounds. Their work should be supported through prioritisation and annotation suggestions based on machine learning technology to efficiently create a large corpus with high-quality annotations for training and evaluating the respective algorithms. Thus, infrastructure should enable the annotation of the same corpus from multiple perspectives by multiple researchers and annotations teams working in parallel. Custom corpora should be importable by the users as needed. Further functionality is needed to maintain and expand the knowledge bases used during the semantic annotation tasks as well as to connect to external standard knowledge bases.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Graphon mean field games with partial observation and application to failure detection in distributed systems
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    MATHIEULOUROCHLAURIERE
  • 依托单位:
基于异构医学影像数据的深度挖掘技术及中枢神经系统重大疾病的精准预测
  • 批准号:
    61672236
  • 项目类别:
    面上项目
  • 资助金额:
    64.0万元
  • 批准年份:
    2016
  • 负责人:
    王骏
  • 依托单位: