NLP4NLP: Applying NLP to Scientific Corpora about Written and Spoken Language Processing

NLP4NLP: Applying NLP to Scientific Corpora about Written and Spoken Language Processing
复制标题

NLP4NLP:将 NLP 应用于有关书面和口语处理的科学语料库

DOI:
--
复制
发表时间:
2015
期刊:
--
影响因子:
--
通讯作者:
P. Paroubek
P. Paroubek
中科院分区:
--
文献类型:
--
作者:
Gil Francopoulo;J. Mariani;P. Paroubek

文献摘要

被引文献

相似文献

分析一个科学领域趋势的演变,以提供关于其状态的见解,并建立关于其未来的可靠假设,是我们在这里要解决的问题。我们通过处理域出版物的元数据和文本内容来解决这个问题。理想情况下,人们希望能够自动综合文档及其元数据中存在的所有信息。作为NLP社区的成员,我们已经将社区开发的工具应用于我们自己领域的出版物,可以称之为“递归”方法。在第一步中,我们收集了NLP会议和期刊的论文语料库,包括文本和语音,涵盖了从60年代到2015年的文档。然后,我们已经挖掘了我们的科学出版物数据库,根据广泛的视角从定量和定性结果绘制了我们领域的图片:从子域,特定社区,年表,术语,概念演变,重复使用和剽窃,趋势预测,新奇检测等等。我们在这里提供了一个帐户的语料库收集和处理与NLP技术,指出每个方面使用的技术。我们得出结论,这样的语料库的域的演员和条件,推广这种方法到其他科学领域带来的好处。会议主题方法和技术,引文和共引分析,科学欺诈和不诚实,自然语言处理
Analyzing the evolutions of the trends of a scientific domain in order to provide insights on its states and to establish reliable hypotheses about its future is the problem we address here. We have approached the problem by processing both the metadata and the text contents of the domain publications. Ideally, one would like to be able to automatically synthesize all the information present in the documents and their metadata. As members of the NLP community, we have applied the tools developed by our community to publications from our own domain, in what could be termed a “recursive” approach. In a first step, we have assembled a corpus of papers from NLP conferences and journals for both text and speech, covering documents produced from the 60’s up to 2015. Then , we have mined our scientific publication database to draw a picture of our field from quantitative and qualitative results according to a wide range of perspectives: ranging from sub-domains, specific communities, chronology, terminology, conceptual evolution, re-use and plagiarism, trend prediction, novelty detection and many more. We provide here an account of the corpus collection and of its processing with NLP technology, indicating for each aspect which technology was used. We conclude on the benefits brought by such corpus to the actors of the domain and on the conditions to generalize this approach to other scientific domains. Conference Topics Methods and techniques, Citation and co-citation analysis, Scientific fraud and dishonesty, Natural Language Processing