课题基金 / 基金详情

Interactive distributed corpus exploration and annotation infrastructure for large corpora and knowledge-bases

Interactive distributed corpus exploration and annotation infrastructure for large corpora and knowledge-bases
适用于大型语料库和知识库的交互式分布式语料库探索和注释基础设施
批准号:
315979217
负责人:
Dr.-Ing. Richard Eckart de Castilho
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research data and software (Scientific Library Services and Information Systems)
财政年份:
2016
资助国家:
德国
项目状态:
已结题
起止时间:
2015-12-31 至 2021-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该项目的目标是建立一个研究基础设施,通过灵活地构建子语料库,将语料库标注扩展到大型文本文档集合。该基础结构满足了计算语言学家和语料库语言学家对通用工具的需求,以便在文档内部和跨文档执行选择性的语义注释任务。这样的基础设施很重要,因为它能够有针对性地利用大量数字可用文本进行语言分析。基础设施应支持专家用户探索大型文档集,建立注释方案,并从大型背景语料库中提取特定于任务的次级语料库。语料库的注释应该可以灵活地分配给不同资质级别和背景的远程工作注释团队。他们的工作应该通过基于机器学习技术的优先顺序和注释建议来支持,以有效地创建一个具有高质量注释的大型语料库,用于训练和评估各自的算法。因此,基础设施应该允许多个研究人员和注释团队并行工作,从多个角度对同一语料库进行注释。自定义语料库应该可以由用户根据需要导入。还需要进一步的功能来维护和扩展在语义注释任务期间使用的知识库以及连接到外部标准知识库。
英文摘要
The goal of this project is a research infrastructure for corpus annotation that scales to large text document collections by flexibly building subcorpora. The infrastructure addresses the needs of computational linguists and corpus linguists for a generic tool to perform selective semantic annotation tasks within and across documents. Such an infrastructure is important because it enables the targeted exploitation of the huge amounts of digitally available text for linguistic analysis. The expert user should be supported by the infrastructure in exploring the large document collections, in setting up an annotation scheme, and in extracting task-specific subcorpora from a large background corpus. The annotation of the corpora should be flexibly distributable to remotely working annotation teams of different qualification levels and backgrounds. Their work should be supported through prioritisation and annotation suggestions based on machine learning technology to efficiently create a large corpus with high-quality annotations for training and evaluating the respective algorithms. Thus, infrastructure should enable the annotation of the same corpus from multiple perspectives by multiple researchers and annotations teams working in parallel. Custom corpora should be importable by the users as needed. Further functionality is needed to maintain and expand the knowledge bases used during the semantic annotation tasks as well as to connect to external standard knowledge bases.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Graphon mean field games with partial observation and application to failure detection in distributed systems
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    MATHIEULOUROCHLAURIERE
  • 依托单位:
基于异构医学影像数据的深度挖掘技术及中枢神经系统重大疾病的精准预测
  • 批准号:
    61672236
  • 项目类别:
    面上项目
  • 资助金额:
    64.0万元
  • 批准年份:
    2016
  • 负责人:
    王骏
  • 依托单位: