Probabilistic knowledge representation of big data for efficient integration and useful inference in bio-medicine
Probabilistic knowledge representation of big data for efficient integration and useful inference in bio-medicine
批准号:
1949035
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Background: In order to realise the opportunities offered by big data in the field of biomedicine,these must be combined, integrated, and made available for cross-cutting research. To achieve this,many issues stemming from the inherent heterogeneity of datasets, such as the use of different datamodels, naming conventions and levels of abstraction, must be overcome. Traditional approachesaimed at improving compatibility of discrete datasets have focused on establishing universalstandards for data representation at the source; for example, by using standard vocabularies orontologies to consistently and unambiguously represent concepts of interest. However, while suchapproaches are ideal in theory, they have had few successful examples. Universal standards arepoorly adopted in the biomedical community. Alternatively, conventional integration methods arelaborious and usually lead to reduction in dimensionality and loss of information in the integrateddataset. We propose the development of a Bayesian approach to harmonisation and dataintegration, allowing for more inclusive and flexible knowledge representation.Hypothesis: Novel methods for data integration, adopting a probabilistic approach, can improve theinference from heterogeneous datasets for biomedical research.Objectives:1. Review existing literature of probabilistic learning/mining methodologies to identify potentialstarting points for this project.2. Extend and/or develop methods for generic data description and visualisation, to assist in rapidexploration of new datasets.3. Develop a Bayesian approach for knowledge representation and integration.4. Evaluate performance, accuracy and limitations, and compare to existing integration methods.5. Apply newly developed methods to real-world datasets in biomedicine and health such ascancer clinical trials and studies in Lupus.Methods: This project requires the development, application, and evaluation of computationalmethods in machine learning, data processing, and knowledge representation. The student will alsobe exposed to methods for data analysis and data mining in the biomedical domain. The Farr@HeRCdeveloped eLab platform, adopted by multiple projects to integrate data gathered acrossinternational studies, will be used as a testbed for the developed methods.Outcomes/Impact:1. A new framework for efficient, probabilistic data integration, with an evaluation of itsperformance, accuracy and limitations.2. New approaches for knowledge representation in the biomedical domain, with potential for wider adoption in other fields.3. One or more applied use cases that demonstrate the utility of the developed methodologies.Training:Dr Niels Peek - Informatics and machine learning input.Dr Nophar Geifman - Informatics, knowledge representation and biomedical input.Dr Philip Couch - Information sciences input.The student will sit in the MRC-funded Farr@HeRC, with exposure to a range of informatics,statistics, clinical, and epidemiological expertise. He/she will belong to the HeRC Doctoral TrainingNetwork and will have the opportunity to receive training as part of this, as well as other in housetraining such as CPD programmes and MSc Health Data Science modules as required
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金