课题基金 / 基金详情

IntelliText - Intelligent Tools for Creating and Analysing Electronic Text Corpora for Humanities Research

IntelliText - Intelligent Tools for Creating and Analysing Electronic Text Corpora for Humanities Research
IntelliText - 用于创建和分析人文研究电子文本语料库的智能工具
批准号:
AH/H037306/1
负责人:
Anthony Hartley
金额:
$20.3万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2010
资助国家:
英国
项目状态:
已结题
起止时间:
2010 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
许多人文科学研究依赖于或将受益于电子语料库的分析--具有代表性的文本集合(如图书、报纸文章、计算机可读格式的技术手册),这些文本也可以用语言或领域信息进行注释。与手工挑选的例子相比,使用语料库的主要优势是能够系统地收集数据,评估某些特征对研究材料的中心性,并在实验中建立数据中的潜在趋势。由于数据分析的一致性增加,依赖电子语料库的项目可望产生更大的学术和社会影响。然而,人文科学研究中基于语料库的研究面临的主要困难是,创建和注释新的语料库和设计用于文本分析的适当搜索引擎需要复杂的技术支持,例如编程、网络开发等方面的专门知识。这样的技术专门知识水平往往无法用于较小的人文项目;但即使是较大的基于语料库的项目也往往因为对相关计算方面的方法或技术支持不足而错过数据分析的机会。即使语料库已经存在,构建适当的计算工具来分析、智能搜索和可视化数据的任务仍然对许多潜在的人文项目来说仍然太具挑战性。人文研究人员缺乏对基于语料库的研究的现代计算技术的认识,这可能严重限制任何计划研究项目的范围和影响。此外,设计基于语料库的工具的计算机科学家往往不了解人文研究的具体需求;他们的工具往往难以适应特定项目,或者缺乏直观的界面和文档。因此,在人文科学中,几种具有收集和准备语料库材料并揭示数据中新的依赖和模式的能力的非平凡计算技术的存在被忽视了。因此,重要的潜在研究协同效应被忽视了。IntelliText的新贡献将是将从计算机科学到人文研究人员的先进工具和方法进行调整,将它们整合到一个具有简单界面和良好文档的单一软件应用程序中。这将允许没有计算机科学或语料库语言学专业背景的人文研究人员利用强大的文本收集和分析方法。它将使他们能够从网络上收集新的项目语料库,用语言学和其他注释自动丰富这些语料库,然后很容易地发现有趣的使用模式,要么从他们自己的直觉和假设开始,要么从系统确定的潜在值得注意的短语和模式开始。研究人员将对翻译文本的文体特征、语言学习和对比语言学以及检测和描述情绪和观点的变化感兴趣,设计软件并在新的应用中进行测试。这些将证明其通用性,以满足范围广泛的人文研究人员的需求,包括历史学家和文学、媒体以及企业或政府传播方面的专家,他们都在项目委员会有代表。IntelliText将作为开源软件免费用于研究目的,将这些工具和方法引入新的领域,并允许用户社区在资助结束后进一步扩展。简而言之,IntelliText的影响将是通过使更多的研究人员能够做出可测试的预测,并通过参考先进和自动化分析技术发现的坚实的语料库证据来加强许多人文学科的理论基础。
英文摘要
Much humanities research relies on or would benefit from analysis of electronic corpora - representative collections of texts (such as books, newspaper articles, technical manuals in computer-readable format), which may also be annotated with linguistic or domain information. The main advantage of using corpora over hand-picked examples is the ability to collect data systematically, to assess the centrality of certain features to the research material, and to establish experimentally potential trends in the data. Projects which rely on electronic corpora can be expected to have greater academic and social impact, thanks to increased consistency in data analysis.However, the major difficulty faced by corpus-based studies in humanities research is that creating and annotating a new corpus and designing an appropriate search engine for textual analysis require complex technical support, e.g., expertise in programming, web development, etc. Such a level of technical expertise is often unavailable to smaller humanities projects; but even larger corpus-based projects often miss opportunities for data analysis because of inadequate methodological or technological support for relevant computational aspects. Even when a corpus already exists, the task of building appropriate computational tools for analysing, intelligently searching and visualising the data still remain too challenging for many potential humanities projects.Humanities researchers' lack of awareness of modern computational techniques for corpus-based studies can seriously limit the scope and the impact of any planned research projects. Moreover, computer scientists who design corpus-based tools frequently do not understand the specific needs of humanities research; their tools are often difficult to adapt to a specific project, or lack an intuitive interface and documentation. As a result, the existence of several non-trivial computational techniques with the power to collect and prepare corpus material and reveal new dependencies and patterns in the data has been overlooked in the humanities. Thus important potential synergies for research have been neglected.IntelliText's novel contribution will be to tune advanced tools and methods from computer science to the needs of humanities researchers, integrating them into a single software application with a simple interface and good documentation. This will allow humanities researchers with no specialised background in computer science or corpus linguistics to take advantage of powerful methods of text collection and analysis. It will enable them to collect new project corpora from the web, have them enriched automatically with linguistic and other annotations, and then easily uncover interesting patterns of usage, starting either from their own intuitions and hypotheses, or from expressions and patterns identified as potentially noteworthy by the system.The software will be designed and tested in novel applications by researchers interested in the stylistic features of translated text, in language learning and contrastive linguistics, and in detecting and describing shifts in sentiment and opinion. These will demonstrate its generalisability for addressing the needs of a wide spectrum of humanities researchers, including historians and specialists in literature, media and corporate or government communications, all of whom are represented on the Project Board. IntelliText will be made freely available for research purposes as Open Source software, introducing these tools and methods into fresh areas and permitting further extensions by the user community after funding ends.In short, the impact of IntelliText will be to strengthen the theoretical foundations of many humanities disciplines by enabling a much larger community of researchers than hitherto to make testable predictions, and then to verify themby reference to solid corpus evidence uncovered by advanced and automated analytical techniques.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Intelligent Patent Analysis for Optimized Technology Stack Selection:Blockchain BusinessRegistry Case Demonstration
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    USHARANI HAREESH GOVINDARA JAN
  • 依托单位: