Formulaic expressions in scientific texts: Corpus design, extraction and exploration

Formulaic expressions in scientific texts: Corpus design, extraction and exploration
复制标题

科学文本中的公式化表达:语料库设计、提取和探索

DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
E. Teich
E. Teich
中科院分区:
--
文献类型:
--
作者:
Hannah Kermes;E. Teich

文献摘要

被引文献

相似文献

Introduction Main Conclusion 229 231 230 226 4408 4687 4639 4801 0.052 0.0493 0.0496 0.0471 Table 7: Type-token-ratio of 4gram formulaic expressions across text parts Figure 7: Ranking of the top 20 formula of Computer Science 114 Hannah Kermes and Elke Teich In contrast to the type-token distribution for academic disciplines, we can observe only very slight differences. The highest density of formulaic expressions is in the Conclusion part, which uses the most formulas, however also the fewest types. The fewest formulas are used in the Abstract. We would have expected a more evident difference. For further exploration, we will have a look at the distribution of structural types of formulas across text parts shown in decreasing order of frequency as a parallel coordinate plot in Figure 10. We can observe that most of the structural types spread more or less evenly throughout the different text parts. The structural type base_VP_mod shows a clear tendency to occur in the Conclusion. The modals can, could are used to refl ect on the presented research, and the modal would is used to express acknowledgments and to point to future work. If we look at the lexical fi llers, the three most frequent formulas in the conclusion are would like to thank, the authors would like, authors would like to, all of which are rather low in frequency in the other text parts. As we look at a specifi c ngram length only, the resulting structures may include formulas which belong to longer units. It is very likely that most of the instances of the three formulas belong to a longer unit ((the authors) would like to thank). For a full coverage, it is thus desirable to include ngrams of different lengths. Figure 8: Comparison of the top 20 formula in A-B1-C1 115 Formulaic expressions in scientifi c texts: Corpus design, extraction and exploration We can further observe a rather high peak for base_VP in the Introduction. In order to explain this, we have to look more closely at the lexical fi llers. There are only 11 different formulas in this class. Three of these formulas occur almost exclusively in the Introduction (around 90% of the occurrences) and are among the top ten of formulas in this text part but not among the top 50 overall: the paper is organized, paper is organized as, is organized as follows. Again, the three 4grams probably belong to a longer unit ((the) paper is organized as (follows)). The picture gets more obvious, if we look at the ranking of the 20 most frequent formulaic expressions occurring in the Introduction (cf. Figure 11). We can see that the ranks of these formulaic expressions are extremely low for all other text parts. All of these formulaic expressions function as markers introducing a specifi c content (e.g. how the paper is structured, which is a typical piece of information in introductory sections). We encounter a similar picture for the other expressions (e.g. of this paper is, in this paper we), which are also used quite frequently in the Abstract and the Conclusion parts as well. Figure 9: Comparison of the top 20 formula in A-B4-C4 116 Hannah Kermes and Elke Teich 5. Conclusions and future work We have presented a methodology for the extraction of formulaic expressions and the calculation of their frequency distributions on an automatic basis. The pipeline we have built for this purpose allows to apply several (related) queries consecutively in order to extract information about the usage of formulas according to selected parameters (here: academic disciplines, text parts). The process is easily reproducible and applicable to other corpora and parameters (provided the necessary information is encoded). The pipeline also includes multiple sorting and grouping options and different kinds of statistical analysis as well as visualization of the results. We have shown selected analyses using the pipeline on a corpus of scientifi c texts, focusing on 4grams. The results show differences with respect to the distribution of formulas and their structural types across academic disciplines and text parts. We could further observe that academic disciplines differ with respect to the density of formulas: Linguistics, Biology and Computational Linguistics are less dense than the other six disciplines included in the corpus. With regard to text parts, the Conclusion has the highest density of formulas, while the other text parts (Abstract, Introduction, Main) are rather similar. To further interpret these results, we need to know about the functions of the formulas we have extracted. We have performed some preliminary experiments in clustering of formulas in order to determine their functions. The results look promising but further information will have to be included in the classifi cation, such as collocation information, syntactic properties (word order changes, syntactic function, etc.) and morpho-syntactic properties (infl ection, determiners, etc.). Figure 10: 4gram distribution of structural types across text parts (frequency per million) 117 Formulaic expressions in scientifi c texts: Corpus design, extraction and exploration Our longer term goals are twofold. First, we want to investigate the diachronic dimension, looking at recent diachronic changes in connection with the evolution of the contact disciplines (i.e. Computational Linguistics, Bioinformatics, Digital Construction, Micro-Electronics). Here, we are interested in processes of diversifi cation (i.e. as a discipline matures, we would expect it to develop distinctive patterns of linguistic variation) as well as standardization (i.e. as a discipline matures, we would expect it to develop a fairly stable set of recurring linguistic patterns with rather little variation). Second, as mentioned in Section 1, we are planning to build a digital resource for use in language pedagogy that may serve both students and teachers as a source of information on scientifi c writing. Essentially this will take the form of an on-line corpus annotated at various linguistic levels and in terms of various linguistic phenomena, including formulaic expressions, very much in the spirit of Davies’ WORD AND PHRASE INFO. The corpus may be queried and/or browsed. We will make additional information available about the annotated formulas, e.g., information about frequency distribution, collocations, structural type, and eventually also functional information. Thus, although not prototypical for electronic dictionaries, the resource can provide valuable lexicographic information. Due to the nature of formulaic expressions, we believe it is essential to have a close connection to a corpus as the usage of formulas is most important for the potential user. Figure 11: Ranking of the 20 most frequent formula in the Introduction 118 Hannah Kermes and Elke Teich We are planning to build a processing pipeline for the annotation process as well. Together with the dedicated processing pipelines for extraction and analysis this will potentially provide the possibility for students and teachers to analyze and annotate their own corpora (either corpora from other registers or student essays). This could potentially reveal shortcomings and strengths of an essay with respect to the usage of formulaic expressions (and possibly other phenomena), and might provide helpful information to better master these important building blocks of discourse. Both the corpus and the processing pipelines will be made available through the CLARIN-D infrastructure.