The Brazilian Portuguese Lexicon: An Instrument for Psycholinguistic Research.

The Brazilian Portuguese Lexicon: An Instrument for Psycholinguistic Research.
复制标题

DOI:
10.1371/journal.pone.0144016
复制
发表时间:
2015
期刊:
影响因子:
3.7
通讯作者:
Meunier F
Meunier F
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Estivalet GL;Meunier F

文献摘要

参考文献

被引文献

相似文献

在这篇文章中,我们提出了巴西葡萄牙语词汇,一个新的基于单词的语料库,在巴西葡萄牙语的心理语言学和计算语言学研究。我们描述了语料库的发展,在互联网上的网站和数据库的用户访问的具体特点。我们还进行分布分析的语料库和比较,以其他当前的数据库。我们的主要目标是提供一个大的,可靠的,有用的基于单词的语料库,具有动态的,易于使用的,直观的界面,可以免费访问互联网的单词和单词标准搜索。我们使用Núcleo Interinformucional de Linguística Computacional的语料库作为基本数据源,并通过导出和添加有关巴西葡萄牙语单词的元语言学和心理语言学信息来开发巴西葡萄牙语词典。我们获得了一个最终的语料库,其中包含超过3000万个单词标记,21.5万个单词类型和每个单词的25类信息。该语料库通过一个免费访问的网站在互联网上提供,该网站有两个搜索引擎:简单搜索和复杂搜索。简单引擎基本上搜索单词列表,而复杂引擎接受语料库类别中的所有类型的标准。输出结果显示了语料库中找到的所有条目,这些条目具有输入搜索中指定的标准,并且可以作为.csv文件下载。我们在结果中创建了一个模块,提供有关每个搜索的基本统计数据。《巴西葡萄牙语词典》还提供了一个伪词引擎以及用于语言和统计分析的具体工具。因此,巴西葡萄牙语词汇是一个方便的工具,刺激搜索,选择,控制和操纵心理语言学实验,因为它也是一个强大的数据库计算语言学研究和语言建模相关的词汇分布,功能和行为。
In this article, we present the Brazilian Portuguese Lexicon, a new word-based corpus for psycholinguistic and computational linguistic research in Brazilian Portuguese. We describe the corpus development, the specific characteristics on the internet site and database for user access. We also perform distributional analyses of the corpus and comparisons to other current databases. Our main objective was to provide a large, reliable, and useful word-based corpus with a dynamic, easy-to-use, and intuitive interface with free internet access for word and word-criteria searches. We used the Núcleo Interinstitucional de Linguística Computacional’s corpus as the basic data source and developed the Brazilian Portuguese Lexicon by deriving and adding metalinguistic and psycholinguistic information about Brazilian Portuguese words. We obtained a final corpus with more than 30 million word tokens, 215 thousand word types and 25 categories of information about each word. This corpus was made available on the internet via a free-access site with two search engines: a simple search and a complex search. The simple engine basically searches for a list of words, while the complex engine accepts all types of criteria in the corpus categories. The output result presents all entries found in the corpus with the criteria specified in the input search and can be downloaded as a.csv file. We created a module in the results that delivers basic statistics about each search. The Brazilian Portuguese Lexicon also provides a pseudoword engine and specific tools for linguistic and statistical analysis. Therefore, the Brazilian Portuguese Lexicon is a convenient instrument for stimulus search, selection, control, and manipulation in psycholinguistic experiments, as also it is a powerful database for computational linguistics research and language modeling related to lexicon distribution, functioning, and behavior.
DOI: 10.3389/fnhum.2015.00004
发表时间: 2015
影响因子: 2.9
作者:
Estivalet GL;Meunier FE
通讯作者: Meunier FE