课题基金 / 基金详情

Linguistic Web Characterization and Web Corpus Construction

Linguistic Web Characterization and Web Corpus Construction
网络语言表征和网络语料库构建
批准号:
261902821
负责人:
Dr. Roland Schäfer
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2014
资助国家:
德国
项目状态:
已结题
起止时间:
2013-12-31 至 2017-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
由网络数据构建的大型语料库通常包含数十亿个标记,特别适合于许多类型的语言研究。因为它们是如此之大,它们允许研究人员研究非常罕见的现象,并且它们包含大量的语言变异。然而,语料库的巨大规模需要通过网络上的无人监督的搜索程序(“爬行”)来收集它们。这通常意味着网络语料库中的文档缺乏大多数语料库语言学家所期望的那种元数据。即使是大型网络语料库的整体构成(就文本类型或注册表而言)也是未知的。此外,迄今为止使用的爬行方法产生了可以证明是有偏见的,即扭曲的样本。最后,Web语料库总是经历完全自动的清理和规范化程序(如删除样板文本元素和去掉重复项),而语料库用户通常对这些程序的精确度和效果几乎一无所知。该项目通过对德国Web进行基本的方法论研究来弥补这种情况。首先,除了传统的有偏爬行过程外,无偏爬行算法被用来生成真正代表Web文档总体的随机样本。此外,为了生成必要的元数据,对文本类型、注册、主题/主题领域等的分类的可用方法被编译并应用于样本。创建资源不是该项目的主要目标,而是开发一种适合于以令人满意的精度应用于大量网络文本的分类方案。为了达到所需的精度,经典的方法(如Biber的多维分析)与信息检索中较新的文档聚类和文档分类方法(如潜在语义分析和主题建模)相结合。一旦元数据可用,大型爬行Web语料库的组成-最重要的也取决于所选的爬行方法-可以第一次指定。此外,通常的清理和标准化程序如何影响网络语料库的组成将是已知的。最后,也是最根本的,用丰富的语言元数据标注的无偏见的代表性Web文档样本的可用性允许对德语Web进行深入的语言特征描述。例如,人们将知道对于一定长度的文件,德语Web的语域构成是什么,等等。这种知识最终使语料库语言学家能够就Web数据和Web语料库是否适合他们的研究问题做出明智的决定。
英文摘要
Large corpora constructed from web data usually contain several billions of tokens and are uniquely suited for many kinds of linguistic research. Because they are so large, they allow researchers to work on very rare phenomena, and they contain a great amount of linguistic variation. However, the immense size of the corpora necessitates their collection by an unsupervised search procedure ("crawling") in the web. This usually means that the documents within web corpora lack the kind of meta data expected by most corpus linguists. Even the overall composition of large web corpora (in terms of text types or registers) is unknown. Furthermore, the crawling methods used so far produce provably biased, i.e., distorted, samples. Finally, web corpora always undergo fully automated cleaning and normalization procedures (such as removal of boilerplate text elements and duplicate removal), and corpus users usually know next to nothing about the precision and the effect of those procedures.This project remedies this situation by performing fundamental methodological research on the German web. Firstly, in addition to conventional biased crawling procedures, unbiased crawling algorithms are used to generate random samples which are truly representative of the population of web documents. Additionally, available methods for the classification of text type, register, subject/topic domain etc. are compiled and applied to the samples in order to generate the necessary meta data. The creation of resources is not the primary goal of this project, but rather the development of a classification scheme which is suitable to be applied to large collections of web texts with satisfactory precision. In order to achieve the required precision, classical methods (e.g. Biber's Multidimensional Analysis) are combined with more recent methods of document clustering and document classification (e.g. Latent Semantic Analysis and Topic Modeling) from Information Retrieval.Once the meta data are available, the composition of large crawled web corpora - most importantly also depending on the chosen crawling method - can be specified for the very first time. Furthermore, it will be known how the usual cleaning and normalization procedures affect the composition of web corpora. Finally and most fundamentally, the availability of unbiased representative web document samples annotated with rich linguistic meta data allows for a deep linguistic characterization of the German web. For example, it will be known what the register composition of the German web is for documents of a certain length, etc. Such knowledge finally puts corpus linguists in a position to make educated decisions as to the suitability of web data and web corpora for their research question.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
On Bias-free Crawling and Representative Web Corpora
无偏差爬行和代表性网络语料库
DOI: 10.18653/v1/w16-2612
发表时间: 2016
期刊:
影响因子: --
作者: [Schäfer, Roland]
通讯作者: Roland
Accurate and efficient general-purpose boilerplate detection for crawled web corpora
用于爬行网络语料库的准确高效的通用样板检测
DOI: 10.1007/s10579-016-9359-2
发表时间: 2017
期刊: Language Resources and Evaluation
影响因子: 2.7
作者: [Schäfer, Roland]
通讯作者: Roland
Abstractions and exemplars: The measure noun phrase alternation in German
摘要和例子:德语中的量词名词短语交替
DOI: 10.1515/cog-2017-0050
发表时间: 2018
期刊: Cognitive Linguistics
影响因子: 1.7
作者: [Schäfer, Roland]
通讯作者: Roland
Automatic Classification by Topic Domain for Meta Data Generation, Web Corpus Evaluation, and Corpus Comparison
按主题域自动分类,用于元数据生成、Web 语料库评估和语料库比较
DOI: 10.18653/v1/w16-2601
发表时间: 2016
期刊:
影响因子: --
作者: [Schäfer, Roland , Felix Bildhauer]
通讯作者: Felix Bildhauer
共 6 条
    国内基金
    海外基金
    基于动态扩散模型与代码知识迁移的Web服务特征增强方法研究
    • 批准号:
      2026JJ80511
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2026
    • 负责人:
      肖勇
    • 依托单位:
    面向Web3D虚拟学习空间的教育智能体系统构建与应用
    • 批准号:
      2025JJ80330
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2025
    • 负责人:
      龙艳军
    • 依托单位:
    基于Web3D元宇宙的实时渲染关键技术研究和应用
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2025
    • 负责人:
      宋三泰
    • 依托单位:
    基于语义理解的多轮多约束Web服务推荐技术
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
    • 依托单位: