课题基金 / 基金详情

PheneBank: automatic extraction and validation of a database of human phenotype-disease associations in the scientific literature

PheneBank: automatic extraction and validation of a database of human phenotype-disease associations in the scientific literature
PheneBank:自动提取和验证科学文献中人类表型与疾病关联的数据库
批准号:
MR/M025160/1
负责人:
Nigel Collier
金额:
$59.12万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2015
资助国家:
英国
项目状态:
已结题
起止时间:
2015 至 --

项目摘要

项目成果

Nigel Collier的其他基金

相似基金

相关文献

中文摘要
翻译
自由文本科学文献有潜力成为揭示基因、疾病和表型之间往往隐藏的关系的极有价值的数据来源。表型描述包括解剖结构、过程和行为的异常。例如“生长迟缓”和“体重减轻”。这些描述构成了确定疾病存在和治疗的基础,但由于其固有的复杂性,以前很少受到文本挖掘社区的关注。近年来,少数专家管理员花费了大量的精力来创建表型编码系统(称为“本体论”),例如人类表型本体(HP)[1]和哺乳动物表型本体(MP)。PheneBank项目建议使用直接从文献中发现的术语来支持和加速策展,并将它们与标准本体自动集成。我们寻求解决三个主要挑战:(1)知识代理:开发最先进的文本挖掘方法来识别科学文本中的表型描述;(2)知识管理:创建科学文本中使用的表型术语的结构化资源,并将其与现有编码系统联系起来;(3)增加对证据的洞察力:与领域专家合作,利用统计关联算法识别有意义的表型-疾病/表型-基因谱。疾病概况将根据人类疾病数据库(例如,人类和孤儿的在线孟德尔遗传)中手工策划的标准进行评估,重点是罕见疾病。挖掘的数据将在一个机器可理解的数据库中提供——这是该项目的明确输出——以支持临床医生和科学家。在技术层面,该项目将开拓利用机器学习(ML)的文本挖掘新方法。由于各种原因,科学文本仍然是一个具有挑战性的领域:描述性命名,高度模糊/词汇外的单词,复杂句子结构的使用以及不断发展的词汇。当前的术语识别技术使用流水线中的机器学习来搜索代表基因、蛋白质和细胞等的连续单词序列。最先进的模型包括使用基于字典的特征集的条件随机场,以及术语所在的本地和主题上下文。然而,表型描述通常由不连续序列表示,例如“患者的生长延迟”。以前没有提到的一个关键方面是捕获这些非规范术语。这需要一种基于语法解析算法的不同范式,这种范式可以捕获结构关系,也需要一种联合学习技术,这种技术可以同时利用大量特征,并在提到表型的不同背景下优化这些特征。该项目还试图利用文本来提取表型、疾病和基因之间具有统计意义的关联。早期的方法因为没有对它们试图瞄准的关系提供深入的语义描述而受挫。这意味着关联评分合并了遗传、药理学和流行病学关系等概念,没有区别。我们基于解析的方法试图通过发现更精确的关系来克服这个问题。该方法遵循了威康信托桑格研究所(WTSI)的突破性工作,包括使用构成表型的概念元素的两两评分来对表型进行术语对齐。该项目的一个令人兴奋的方面是在利益相关者之间进行跨学科合作,以建立表型疾病概况资源:(a)来自剑桥大学、科罗拉多大学和曼彻斯特大学的计算机科学家;(b)来自WTSI、麦吉尔大学和EMBL-EBI的生物信息学家和生命科学家,以及(c)来自NIHR Bioresource的临床医生。
英文摘要
Free text scientific literature has the potential to be an incredibly valuable source of data for uncovering the often hidden relationships between genes, diseases and phenotypes. Phenotypic descriptions cover abnormalities in anatomical structures, processes and behaviours. For example 'growth delay' and 'body weight loss'. Such descriptions form the basis for determining the existence and treatment of a disease but, because of their inherent complexity, have previously received less attention by the text mining community. In recent years, significant effort has been spent by a small number of expert curators to create coding systems for phenotypes (called "ontologies"), such as the Human Phenotype Ontology (HP) [1] and the Mammalian Phenotype Ontology (MP). The PheneBank project proposes to support and speed up curation using terms discovered directly from the literature and to automatically integrate them with such standard ontologiesThere are three major challenges we seek to address: (1) knowledge brokering: to develop state of the art text mining approaches to identify phenotypic descriptions in scientific texts; (2) knowledge management: to create a structured resource of phenotype terms used in scientific texts and link them to existing coding systems; and (3) adding insight to evidence: to work with domain experts to utilize statistical association algorithms to identify meaningful phenotype-disease / phenotype-gene profiles. The disease profiles will be evaluated against hand curated standards in human disease databases (e.g. Online Mendelian Inheritance of Man and OrphaNet) with a focus on rare diseases. Mined data will be provided in a machine understandable database - a definitive output of the project - to support clinicians and scientists. At the technological level the project will pioneer new methods for text mining that exploit machine learning (ML). Scientific texts remain a challenging area for a variety of reasons: descriptive naming, high levels of ambiguity/out of vocabulary words, use of complex sentence structures and an evolving vocabulary. Current techniques in term recognition employ ML in pipelines to search for continuous sequences of words that represent genes, proteins and cells etc. State of the art models include conditional random fields using feature sets based on dictionaries as well as the local and topical context where the term is located. However, phenotype descriptions are often represented by discontinuous sequences, such as 'growth in the patient was delayed'. One key aspect not previously addressed is in the capture of such non-canonical terms. This requires a different paradigm based on grammatical parsing algorithms that capture structural relations as well as joint learning techniques that can leverage large numbers of features simultaneously and optimise these across the diverse contexts in which phenotypes are mentioned.The project also seeks to harness texts for extracting statistically significant associations between phenotypes, diseases and genes. Earlier approaches have suffered from not providing deep semantic descriptions of the relations they tried to target. This means that association scores merge notions of genetic, pharmacological, and epidemiological relations etc. without distinction. Our parsing-based approach is an attempt to overcome this issue by discovering more precise relationships. The approach follows ground breaking work at the Wellcome Trust Sanger Institute (WTSI), including terminology alignment of phenotypes using pairwise scoring of the conceptual elements that make up the phenotype. An exciting aspect of this project is inter-disciplinary collaboration across stakeholders to build a resource of phenotype-disease profiles: (a) computer scientists from the Universities of Cambridge, Colorado and Manchester; (b) bioinformaticians and life scientists from the WTSI, McGill University and EMBL-EBI, and (c) clinicians from the NIHR Bioresource.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.18653/v1/d18-1250
发表时间: 2018
期刊:
影响因子: --
作者: [Hoang-Quynh Le;Duy-Cat Can;Sinh T. Vu;T. Dang;Mohammad Taher Pilehvar;Nigel Collier]
通讯作者: Hoang-Quynh Le;Duy-Cat Can;Sinh T. Vu;T. Dang;Mohammad Taher Pilehvar;Nigel Collier
A pragmatic guide to geoparsing evaluation
地理解析评估实用指南
DOI: 10.17863/cam.55940
发表时间: 2019
期刊:
影响因子: --
作者: [Gritta M]
通讯作者: Gritta M
EPI-AI: Automated Understanding and Alerting of Disease Outbreaks from Global News Media
  • 批准号:
    ES/T012277/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $62.61万
  • 财政年份:
    2020
  • 负责人:
    Nigel Collier
  • 依托单位:
SIPHS: Semantic interpretation of personal health messages for generating public health summaries
  • 批准号:
    EP/M005089/1
  • 项目类别:
    Fellowship
  • 资助金额:
    $123.85万
  • 财政年份:
    2015
  • 负责人:
    Nigel Collier
  • 依托单位:
国内基金
海外基金
基于计算模型的医用X线最优曝光控制技术的研究
  • 批准号:
    60472004
  • 项目类别:
    面上项目
  • 资助金额:
    26.0万元
  • 批准年份:
    2004
  • 负责人:
    牟轩沁
  • 依托单位: