Linguistic Web Characterization and Web Corpus Construction
Linguistic Web Characterization and Web Corpus Construction
批准号:
261902821
负责人:
Dr. Roland Schäfer
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2014
资助国家:
德国
项目状态:
已结题
起止时间:
2013-12-31 至 2017-12-31
中文摘要
由网络数据构建的大型语料库通常包含数十亿个标记,并且特别适合于多种语言学研究。因为它们是如此之大,它们允许研究人员研究非常罕见的现象,它们包含了大量的语言变异。然而,语料库的巨大规模需要他们的收集在网络中的无监督搜索过程(“爬行”)。这通常意味着网络语料库中的文档缺乏大多数语料库语言学家所期望的那种Meta数据。甚至大型网络语料库的整体组成(就文本类型或寄存器而言)也是未知的。此外,到目前为止使用的爬行方法产生可证明的偏差,即,扭曲的样本最后,网络语料库总是经历完全自动化的清理和规范化过程(如删除样板文本元素和重复删除),语料库用户通常对这些过程的精度和效果一无所知。首先,除了传统的有偏爬行程序,无偏爬行算法被用来生成随机样本,这是真正代表的人口的Web文档。此外,可用的方法进行分类的文本类型,寄存器,主题/主题域等编译和应用到样本,以生成必要的Meta数据。资源的创建不是这个项目的主要目标,而是开发一个分类方案,适用于大量的Web文本集合,并具有令人满意的精度。为了达到所需的精度,经典方法(例如Biber的多维分析)与最近的文档聚类和文档分类方法相结合(例如潜在语义分析和主题建模)。一旦Meta数据可用,可以第一次指定大型爬行网络语料库的组成--最重要的是也取决于所选择的爬行方法。此外,它将是已知的通常的清洗和规范化程序如何影响网络语料库的组成。最后,也是最根本的,无偏见的代表性的Web文档样本注释丰富的语言Meta数据的可用性允许德国网络的深层语言特征。例如,它将是已知的德语网络的寄存器组成是为一定长度的文件等,这样的知识最终使语料库语言学家能够作出明智的决定,以适合他们的研究问题的网络数据和网络语料库。
英文摘要
Large corpora constructed from web data usually contain several billions of tokens and are uniquely suited for many kinds of linguistic research. Because they are so large, they allow researchers to work on very rare phenomena, and they contain a great amount of linguistic variation. However, the immense size of the corpora necessitates their collection by an unsupervised search procedure ("crawling") in the web. This usually means that the documents within web corpora lack the kind of meta data expected by most corpus linguists. Even the overall composition of large web corpora (in terms of text types or registers) is unknown. Furthermore, the crawling methods used so far produce provably biased, i.e., distorted, samples. Finally, web corpora always undergo fully automated cleaning and normalization procedures (such as removal of boilerplate text elements and duplicate removal), and corpus users usually know next to nothing about the precision and the effect of those procedures.This project remedies this situation by performing fundamental methodological research on the German web. Firstly, in addition to conventional biased crawling procedures, unbiased crawling algorithms are used to generate random samples which are truly representative of the population of web documents. Additionally, available methods for the classification of text type, register, subject/topic domain etc. are compiled and applied to the samples in order to generate the necessary meta data. The creation of resources is not the primary goal of this project, but rather the development of a classification scheme which is suitable to be applied to large collections of web texts with satisfactory precision. In order to achieve the required precision, classical methods (e.g. Biber's Multidimensional Analysis) are combined with more recent methods of document clustering and document classification (e.g. Latent Semantic Analysis and Topic Modeling) from Information Retrieval.Once the meta data are available, the composition of large crawled web corpora - most importantly also depending on the chosen crawling method - can be specified for the very first time. Furthermore, it will be known how the usual cleaning and normalization procedures affect the composition of web corpora. Finally and most fundamentally, the availability of unbiased representative web document samples annotated with rich linguistic meta data allows for a deep linguistic characterization of the German web. For example, it will be known what the register composition of the German web is for documents of a certain length, etc. Such knowledge finally puts corpus linguists in a position to make educated decisions as to the suitability of web data and web corpora for their research question.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
On Bias-free Crawling and Representative Web Corpora
无偏差爬行和代表性网络语料库
DOI:
10.18653/v1/w16-2612
发表时间:
2016
期刊:
影响因子:
--
作者:
[Schäfer, Roland]
通讯作者:
Roland
Accurate and efficient general-purpose boilerplate detection for crawled web corpora
用于爬行网络语料库的准确高效的通用样板检测
DOI:
10.1007/s10579-016-9359-2
发表时间:
2017
期刊:
Language Resources and Evaluation
影响因子:
2.7
作者:
[Schäfer, Roland]
通讯作者:
Roland
Abstractions and exemplars: The measure noun phrase alternation in German
摘要和例子:德语中的量词名词短语交替
DOI:
10.1515/cog-2017-0050
发表时间:
2018
期刊:
Cognitive Linguistics
影响因子:
1.7
作者:
[Schäfer, Roland]
通讯作者:
Roland
Automatic Classification by Topic Domain for Meta Data Generation, Web Corpus Evaluation, and Corpus Comparison
按主题域自动分类,用于元数据生成、Web 语料库评估和语料库比较
DOI:
10.18653/v1/w16-2601
发表时间:
2016
期刊:
影响因子:
--
作者:
[Schäfer, Roland , Felix Bildhauer]
通讯作者:
Felix Bildhauer
DOI:
10.1515/9783110518214-020
发表时间:
2016
期刊:
影响因子:
--
作者:
[Bildhauer, Felix , Roland Schäfer]
通讯作者:
Roland Schäfer
共 6 条
国内基金
海外基金
登录
查看更多内容
基于动态扩散模型与代码知识迁移的Web服务特征增强方法研究
-
批准号:2026JJ80511
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:肖勇
-
依托单位:
面向Web3D虚拟学习空间的教育智能体系统构建与应用
-
批准号:2025JJ80330
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:龙艳军
-
依托单位:
基于Web3D元宇宙的实时渲染关键技术研究和应用
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:宋三泰
-
依托单位:
基于语义理解的多轮多约束Web服务推荐技术
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
Web大数据环境下基于迁移学习的跨领域推荐研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
数据智能驱动的泛在Web应用服务质量优化方法研究
-
批准号:62102009
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:马郓
-
依托单位:
基于侧信道分析的Web站点指纹识别技术研究
-
批准号:62102084
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:顾晓丹
-
依托单位:
基于时间意图的地表覆盖Web 信息发现方法研究
-
批准号:2021JJ40721
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2021
-
负责人:侯东阳
-
依托单位:
恶劣条件下Web服务QoS预测与QoS确保的服务组合卸载方法研究
-
批准号:--
-
项目类别:面上项目
-
资助金额:58万元
-
批准年份:2021
-
负责人:夏云霓
-
依托单位:
多模态Web信息检索排序学习方法研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2021
-
负责人:耿光刚
-
依托单位: