Lexical Acquisition for the Biomedical Domain
Lexical Acquisition for the Biomedical Domain
批准号:
EP/G051070/1
负责人:
Anna Korhonen
金额:
$36.37万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2009
资助国家:
英国
项目状态:
已结题
起止时间:
2009 至 --
中文摘要
现在迫切需要自然语言处理(NLP)来协助处理、挖掘和提取生物医学领域快速增长的文献中的知识。近年来,生物医学基础自然语言处理技术的发展取得了长足的进展。当前的挑战是通过更丰富和更深入的分析来改进这些技术,从而能够支持广泛的现实世界任务。为此,高质量的词汇资源(如准确和全面的词汇和词分类)是非常需要的。当前系统中使用的大多数词汇资源都是由语言学家手工开发的。手工工作是非常昂贵的,并且产生的资源需要大量的劳动密集型移植到新的(子)领域和任务。从未注释的文本库(如生物医学文章的语料库)中自动获取或更新词汇信息是一个更有前途的途径。由于词汇习得直接从相关数据中收集使用和频率信息,因此可以大大提高自然语言处理技术的可行性和可移植性。对自动词汇习得的研究现在开始为实际的自然语言处理任务提供大量有用的资源。然而,这些技术在生物医学文本中的应用受到限制,因为许多现有技术需要适应才能在这个具有语言学挑战性的领域中发挥最佳作用。在这个项目中,我们将采用能够从语料库数据中获取动词基本语法语义信息的现有技术,并将其应用于生物医学领域。我们将重点关注口头(i)子分类框架,(ii)选择偏好,以及(ii)词汇语义类。这些信息,当针对问题领域进行定制时,可以帮助关键的NLP任务,如解析、回指解析、信息提取(IE)和问答(QA)。在我们的试点研究和扩展自适应的,最先进的文本处理工具的基础上,我们将进一步改进现有的技术,并使用能够支持有效领域适应的新颖的无监督和半监督方法对其进行扩展。我们将直接在实际的生物-自然语言处理任务的背景下评估和展示我们的技术的能力。我们将使用系统的最终版本从生物医学语料库中获取大量词汇数据库。由此产生的资源将与软件一起免费分发给研究界,该软件可用于将存储在数据库中的频率信息调整为特定的生物医学子领域/任务。我们希望这个项目能够(i)推进生物-自然语言处理并提高其在生物医学实际任务中的实用性,(ii)通过提高词汇习得的准确性、鲁棒性和可移植性来推进自然语言处理,以及(iii)在词汇习得的关键领域提供重要的大规模领域适应研究。
英文摘要
Natural Language Processing (NLP) is now critically needed to assist the processing, mining and extraction of knowledge from the rapidly growing literature in the area of biomedicine. In recent years, considerable progress has been made in the development of basic NLP techniques for biomedicine. The current challenge is to improve these techniques with richer and deeper analysis capable of supporting a wide range of real-world tasks. High-quality lexical resources (e.g. accurate and comprehensive lexicons and word classifications) are critically needed for this. Most lexical resources used in current systems are developed manually by linguists. Manual work is extremely costly, and the resulting resources require extensive labour-intensive porting to new (sub-)domains and tasks. Automatic acquisition or updating of lexical information from repositories of un-annotated text (e.g. corpora of biomedical articles) is a more promising avenue to pursue. Since lexical acquisition gathers usage and frequency information directly from relevant data, it can considerably enhance the viability and portability of NLP technology. Research into automatic lexical acquisition is now starting to produce large-scale resources useful for practical NLP tasks. However, the application of such techniques to biomedical texts has been limited because many existing techniques require adaptation before they can perform optimally in this linguistically challenging domain. In this project, we will take existing techniques capable of acquiring basic syntactic-semantic information for verbs from corpus data and will adapt them to the biomedical domain. We will focus on verbal (i) subcategorization frames, (ii) selectional preferences, and (ii) lexical-semantic classes. This information, when tailored to the domain in question, can aid key NLP tasks such as parsing, anaphora resolution, Information Extraction (IE), and question-answering (QA). Building on our pilot studies and expanding on the adaptive, state-of-the-art text processing tools available to us, we will improve existing techniques further and extend them with novel unsupervised and semi-supervised methods capable of supporting efficient domain adaptation. We will evaluate and demonstrate the capabilities of our techniques directly and in the context of practical BIO-NLP tasks. We will use the final version of the system to acquire a substantial lexical database from a biomedical corpus. The resulting resource will be distributed freely to the research community, along with the software which can be used to tune the frequency information stored in the database to particular biomedical sub-domains/tasks.We expect this project to (i) advance BIO-NLP and improve its usefulness for practical tasks in biomedicine, (ii) advance NLP by improving the accuracy, robustness and portability of lexical acquisition to real-world tasks, and (iii) provide an important large-scale study of domain-adaptation in the critical area of lexical acquisition.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2012-12
期刊:
影响因子:
--
作者:
[Danish Contractor;Yufan Guo;A. Korhonen]
通讯作者:
Danish Contractor;Yufan Guo;A. Korhonen
DOI:
--
发表时间:
2010-07
期刊:
Laser Physics
影响因子:
1.2
作者:
[Diarmuid Ó Séaghdha]
通讯作者:
Diarmuid Ó Séaghdha
DOI:
10.1186/1471-2105-12-69
发表时间:
2011-03-08
期刊:
BMC bioinformatics
影响因子:
3
作者:
[Guo Y, Korhonen A, Liakata M, Silins I, Hogberg J, Stenius U]
通讯作者:
Stenius U
DOI:
10.1093/bioinformatics/btt163
发表时间:
2013-06
期刊:
Bioinformatics
影响因子:
5.8
作者:
[Yufan Guo;Ilona Silins;U. Stenius;A. Korhonen]
通讯作者:
Yufan Guo;Ilona Silins;U. Stenius;A. Korhonen
DOI:
10.1186/1471-2105-12-212
发表时间:
2011-05-27
期刊:
BMC bioinformatics
影响因子:
3
作者:
[Lippincott T, Séaghdha DÓ, Korhonen A]
通讯作者:
Korhonen A
共 9 条
Towards Globally Equitable Language Technologies (EQUATE)
-
批准号:EP/Y031350/1
-
项目类别:Research Grant
-
资助金额:$269.65万
-
财政年份:2023
-
负责人:Anna Korhonen
-
依托单位:
Literature-based discovery for cancer biology
-
批准号:MR/M013049/1
-
项目类别:Research Grant
-
资助金额:$50.37万
-
财政年份:2015
-
负责人:Anna Korhonen
-
依托单位:
Using Text Mining to Aid Cancer Risk Assessment
-
批准号:G0601766/1
-
项目类别:Research Grant
-
资助金额:$10.88万
-
财政年份:2007
-
负责人:Anna Korhonen
-
依托单位:
海外基金