Automated Biological Event Extraction from the Literature for Drug Discovery
Automated Biological Event Extraction from the Literature for Drug Discovery
批准号:
BB/G013160/1
负责人:
Sophia Ananiadou
金额:
$36.76万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2009
资助国家:
英国
项目状态:
已结题
起止时间:
2009 至 --
中文摘要
新药的开发既昂贵又耗时:即使我们在生命科学领域取得了许多进展,一种新药也可能需要十多年的时间才能被证明是有效和安全的。从一批有希望的早期候选人中,只有少数人最终会被批准。候选人在被发现无法使用(自然减员)之前持续的时间越长,成本就越昂贵,特别是如果涉及临床试验的话。损耗率约为90%,因此损耗对制药行业来说是毁灭性的代价,因此迫切需要减少其影响。英国研究人员在生物和制药研究方面处于领先地位,他们将从尽早识别可能失败的候选药物的方法中受益匪浅,最好是在达到临床阶段之前很久。目前另一个令人关切的领域是药物如何针对群体:并非每个人对同一种药物的反应都是一样的。如果我们能发现哪些基因与此有关,那么我们就可以希望既能专注于更有前途的候选药物,又能找到为个体(群体)量身定制治疗的方法。然而,不幸的是,科学家们面临着严重的知识差距:没有科学家能够使用传统手段跟上生命科学中正在(和已经)产生的大量实验数据,特别是大量相关文献。此外,许多知识隐藏在文献中:已经表明,在文献中可以发现全新的知识,通常是多年,但文献的浩瀚使研究人员无法达到所需的信息检索水平,这是将其链接和合成为新的,以前未知的知识的第一步。信息发现的主要目标是MEDLINE资源,目前包含约1700万摘要:这似乎很大,但仍然是相关全文科学文章中包含的信息和隐藏知识的一小部分。拟议的项目旨在通过开发自动过滤信息和从科学文献中综合新知识的方法,帮助科学家克服这一知识差距。由于蛋白质与生理或病理过程之间的直接联系并不总是在文本中明确描述,我们必须寻找间接证据。这涉及寻找与蛋白质相关的生物过程的迹象。在写作时,生物学家基本上描述了“事件”,例如参与更高级生物过程(如血管生成)的磷酸化。通过识别和提取这些事件,以及特定的生物实体(蛋白质,疾病),我们可以从成千上万的文本中收集有关生物过程的许多信息片段。然后,这些片段可以通过在片段之间建立关联来发现新的知识。为了实现这样的提取片段的知识发现,强大的语义文本挖掘技术,需要能够处理生物学家的特殊语言,并可以实现适当的抽象层次远远超出单纯的单词搜索。该项目将定制国家文本挖掘中心的通用工具,并开展研究,以找到从文献中提取有关生物过程事件的最佳方法。阿斯利康将密切参与,包括为研究提供信息,并提供实用的领域专业知识,要求,数据和具体的评估方案。他们的兴趣还表现在为该项目提供大量现金捐助。该计划的结果将是为学术研究人员提供的文本挖掘服务,提供NaCTeM,支持他们从文献中发现蛋白质-生物过程关联的任务。
英文摘要
The development of new drugs is both expensive and time-consuming: it can take over a decade for a new drug to be proven effective and safe, even with the many advances we have seen in the life sciences. From a batch of promising early candidates, only a few will eventually be approved. The longer a candidate lasts before being found unusable (attrition), the more expensive the cost, especially if clinical trials have been involved. Attrition rates run at ca 90%, and attrition is thus ruinously costly to the pharmaceutical industry, so there is an urgent need to reduce its impact. UK researchers, leading in biological and pharmaceutical research, would benefit greatly from means to identify as early as possible drug candidates that are likely to fail, preferably long before the clinical stage is reached. Another current area of concern is how drugs may be targeted to groups of individuals: not every individual responds in the same way to the same drug.. If we can discover which genes are implicated in this, then we can hope both to focus on the more promising drug candidates and find ways of tailoring treatments to (groups of) individuals. Unfortunately, however, scientists are faced with a severe knowledge gap: no scientist can keep up, using traditional means, with the vast amount of experimental data and especially its massive associated literature that is being (and has been )generated in the life sciences. Moreover, much knowledge is hidden in the literature: it has been shown that entirely new knowledge has been available for discovery in the literature, often for many years, but that the vastness of the literature has prevented researchers from achieving the required level of information retrieval, that is the first step in linking and synthesizing it into new, previously unsuspected knowledge. The main target of information finding is the MEDLINE resource, which currently contains some 17 million abstracts: this is seemingly large but is nevertheless a fraction of the information and hidden knowledge contained in the associated full text scientific articles. The proposed project is designed to help scientists overcome this knowledge gap, by developing automatic means to filter information and to synthesise new knowledge from the scientific literature. As a direct link between a (number of) proteins(s) and a physiological or pathophysiological process is not always described explicitly in a text, we must hunt for indirect evidence. This involves looking for indications of biological processes that are associated with proteins. When writing, biologists essentially describe 'events' such as such as phosphorylation that are involved in higher order bioprocesses such as angiogenesis. By identifying and extracting such events, and the particular biological entities (proteins, diseases), we can collect many fragments of information about bioprocesses from many thousands of texts. These fragments can then be used to find new knowledge by establishing associations among the fragments. To achieve such extraction of fragments for knowledge finding, powerful semantic text mining techniques are required that can handle the special languages of biologists, and that can achieve appropriate levels of abstraction far beyond mere word search. This project will customise the generic tools of the National Centre for Text Mining and carry out research to find the best ways of extracting events concerning biological processes from the literature. AstraZeneca will be closely involved, both in terms of informing the research, and providing practical domain expertise, requirements, data and concrete evaluation scenarios. Their interest is also manifest in a substantial cash contribution to the project. The result of this programme will be a text mining service to academic researchers, offered NaCTeM, supporting them in their task of discovering protein -bioprocess associations from the literature.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1371/journal.pone.0014780
发表时间:
2011-03-29
期刊:
PloS one
影响因子:
3.7
作者:
[Ananiadou S, Sullivan D, Black W, Levow GA, Gillespie JJ, Mao C, Pyysalo S, Kolluru B, Tsujii J, Sobral B]
通讯作者:
Sobral B
Adding text mining workflows as web services to the BioCatalogue
将文本挖掘工作流程作为 Web 服务添加到 BioCatalogue
DOI:
10.1145/2166896.2166913
发表时间:
2011
期刊:
影响因子:
--
作者:
[Kontonasios G]
通讯作者:
Kontonasios G
DOI:
10.1186/1471-2105-14-2
发表时间:
2013-01-16
期刊:
BMC bioinformatics
影响因子:
3
作者:
[Mihăilă C, Ohta T, Pyysalo S, Ananiadou S]
通讯作者:
Ananiadou S
DOI:
10.2196/26892
发表时间:
2021-06-15
期刊:
Journal of medical Internet research
影响因子:
7.4
作者:
[Deng L, Chen L, Yang T, Liu M, Li S, Jiang T]
通讯作者:
Jiang T
DOI:
10.1186/1471-2105-12-s8-s3
发表时间:
2011-10-03
期刊:
BMC bioinformatics
影响因子:
3
作者:
[Krallinger M, Vazquez M, Leitner F, Salgado D, Chatr-Aryamontri A, Winter A, Perfetto L, Briganti L, Licata L, Iannuccelli M, Castagnoli L, Cesareni G, Tyers M, Schneider G, Rinaldi F, Leaman R, Gonzalez G, Matos S, Kim S, Wilbur WJ, Rocha L, Shatkay H, Tendulkar AV, Agarwal S, Liu F, Wang X, Rak R, Noto K, Elkan C, Lu Z, Dogan RI, Fontaine JF, Andrade-Navarro MA, Valencia A]
通讯作者:
Valencia A
共 8 条
Japan Partnering Award. Text mining and bioinformatics platforms for metabolic pathway modelling.
-
批准号:BB/P025684/1
-
项目类别:Research Grant
-
资助金额:$5.07万
-
财政年份:2017
-
负责人:Sophia Ananiadou
-
依托单位:
Enriching Metabolic PATHwaY models with evidence from the literature (EMPATHY)
-
批准号:BB/M006891/1
-
项目类别:Research Grant
-
资助金额:$75.68万
-
财政年份:2015
-
负责人:Sophia Ananiadou
-
依托单位:
Supporting Evidence-based Public Health Interventions using Text Mining
-
批准号:MR/L01078X/1
-
项目类别:Research Grant
-
资助金额:$83.55万
-
财政年份:2014
-
负责人:Sophia Ananiadou
-
依托单位:
Mining the History of Medicine
-
批准号:AH/L00982X/1
-
项目类别:Research Grant
-
资助金额:$33.31万
-
财政年份:2014
-
负责人:Sophia Ananiadou
-
依托单位:
From text to pathways: text mining techniques for reconstructing signalling pathways
-
批准号:BB/G53025X/1
-
项目类别:Research Grant
-
资助金额:$4.34万
-
财政年份:2009
-
负责人:Sophia Ananiadou
-
依托单位:
Tools for the text mining-based visualisation of the provenance of biochemical networks
-
批准号:BB/E004431/1
-
项目类别:Research Grant
-
资助金额:$70.01万
-
财政年份:2007
-
负责人:Sophia Ananiadou
-
依托单位:
海外基金