PheneBank: automatic extraction and validation of a database of human phenotype-disease associations in the scientific literature
PheneBank: automatic extraction and validation of a database of human phenotype-disease associations in the scientific literature
批准号:
MR/M025160/1
负责人:
Nigel Collier
金额:
$59.12万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2015
资助国家:
英国
项目状态:
已结题
起止时间:
2015 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Free text scientific literature has the potential to be an incredibly valuable source of data for uncovering the often hidden relationships between genes, diseases and phenotypes. Phenotypic descriptions cover abnormalities in anatomical structures, processes and behaviours. For example 'growth delay' and 'body weight loss'. Such descriptions form the basis for determining the existence and treatment of a disease but, because of their inherent complexity, have previously received less attention by the text mining community. In recent years, significant effort has been spent by a small number of expert curators to create coding systems for phenotypes (called "ontologies"), such as the Human Phenotype Ontology (HP) [1] and the Mammalian Phenotype Ontology (MP). The PheneBank project proposes to support and speed up curation using terms discovered directly from the literature and to automatically integrate them with such standard ontologiesThere are three major challenges we seek to address: (1) knowledge brokering: to develop state of the art text mining approaches to identify phenotypic descriptions in scientific texts; (2) knowledge management: to create a structured resource of phenotype terms used in scientific texts and link them to existing coding systems; and (3) adding insight to evidence: to work with domain experts to utilize statistical association algorithms to identify meaningful phenotype-disease / phenotype-gene profiles. The disease profiles will be evaluated against hand curated standards in human disease databases (e.g. Online Mendelian Inheritance of Man and OrphaNet) with a focus on rare diseases. Mined data will be provided in a machine understandable database - a definitive output of the project - to support clinicians and scientists. At the technological level the project will pioneer new methods for text mining that exploit machine learning (ML). Scientific texts remain a challenging area for a variety of reasons: descriptive naming, high levels of ambiguity/out of vocabulary words, use of complex sentence structures and an evolving vocabulary. Current techniques in term recognition employ ML in pipelines to search for continuous sequences of words that represent genes, proteins and cells etc. State of the art models include conditional random fields using feature sets based on dictionaries as well as the local and topical context where the term is located. However, phenotype descriptions are often represented by discontinuous sequences, such as 'growth in the patient was delayed'. One key aspect not previously addressed is in the capture of such non-canonical terms. This requires a different paradigm based on grammatical parsing algorithms that capture structural relations as well as joint learning techniques that can leverage large numbers of features simultaneously and optimise these across the diverse contexts in which phenotypes are mentioned.The project also seeks to harness texts for extracting statistically significant associations between phenotypes, diseases and genes. Earlier approaches have suffered from not providing deep semantic descriptions of the relations they tried to target. This means that association scores merge notions of genetic, pharmacological, and epidemiological relations etc. without distinction. Our parsing-based approach is an attempt to overcome this issue by discovering more precise relationships. The approach follows ground breaking work at the Wellcome Trust Sanger Institute (WTSI), including terminology alignment of phenotypes using pairwise scoring of the conceptual elements that make up the phenotype. An exciting aspect of this project is inter-disciplinary collaboration across stakeholders to build a resource of phenotype-disease profiles: (a) computer scientists from the Universities of Cambridge, Colorado and Manchester; (b) bioinformaticians and life scientists from the WTSI, McGill University and EMBL-EBI, and (c) clinicians from the NIHR Bioresource.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI:
10.18653/v1/d18-1250
发表时间:
2018
期刊:
影响因子:
--
作者:
[Hoang-Quynh Le;Duy-Cat Can;Sinh T. Vu;T. Dang;Mohammad Taher Pilehvar;Nigel Collier]
通讯作者:
Hoang-Quynh Le;Duy-Cat Can;Sinh T. Vu;T. Dang;Mohammad Taher Pilehvar;Nigel Collier
A pragmatic guide to geoparsing evaluation
地理解析评估实用指南
DOI:
10.17863/cam.55940
发表时间:
2019
期刊:
影响因子:
--
作者:
[Gritta M]
通讯作者:
Gritta M
EPI-AI: Automated Understanding and Alerting of Disease Outbreaks from Global News Media
-
批准号:ES/T012277/1
-
项目类别:Research Grant
-
资助金额:$62.61万
-
财政年份:2020
-
负责人:Nigel Collier
-
依托单位:
SIPHS: Semantic interpretation of personal health messages for generating public health summaries
-
批准号:EP/M005089/1
-
项目类别:Fellowship
-
资助金额:$123.85万
-
财政年份:2015
-
负责人:Nigel Collier
-
依托单位:
国内基金
海外基金
基于计算模型的医用X线最优曝光控制技术的研究
-
批准号:60472004
-
项目类别:面上项目
-
资助金额:26.0万元
-
批准年份:2004
-
负责人:牟轩沁
-
依托单位: