Using semantics to leverage health and research Big Data
Using semantics to leverage health and research Big Data
批准号:
MR/S003703/1
负责人:
Tim Beck
金额:
$37.35万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2018
资助国家:
英国
项目状态:
已结题
起止时间:
2018 至 --
中文摘要
通过连接与健康相关的研究大数据,可以对它们进行比较,并发现更大的样本/参与者规模以进行分析。大数据的复杂性给集成带来了挑战。表型是一种数据类型,它显示了临床和非临床(例如“组学”)健康研究大数据的特定多样性和变异性。术语“表型”用于定义一组医学上和语义上截然不同的概念,如特征(如血糖水平)、医学体征和症状(如高血糖)和疾病(如2型糖尿病)。为了能够比较大数据中的值,我们需要知道这些值在数据集之间是否具有相同的含义,而执行此操作所需的语义严谨性是通过使用“本体”来提供的。有几个本体描述重叠的表型域,但已被开发用于不同的目的,例如NHS使用SNOMED CT和研究数据库使用人类表型本体(HPO)。如果数据集被编码为不同的本体,或者根本没有本体,那么它们可以通过协调过程链接到共同的本体。这包括映射来自不同本体的术语,以及自由文本的文本挖掘(TM)以将原始值与本体代码相关联。不幸的是,这其中有几个障碍,例如本体和公开可用的本体映射之间的差距,需要调整当前的TM方法,这些方法经过优化以在明确定义的领域中表现良好,以及需要扩展当前的映射和TM方法来处理大数据。在此期间,我将通过使用本体来协调表型数据,创建连接临床和非临床研究大数据的增强能力。这将涉及弥合当前本体中的差距,并采用当前最先进的TM和本体映射方法,以便它们针对这一背景进行优化,并可应用于大数据。我开发的方法将与疾病无关,但首先它们将应用于当地感兴趣的疾病领域。莱斯特生物医学研究中心专注于心血管、呼吸和生活方式疾病,包括初级和二级保健的数据集,以及包括参与者问卷和生物样本数据在内的临床研究研究。我将连接疾病区域内和疾病区域之间的临床研究数据,以及与当地和公共可用的组学数据,例如基因组和表观基因组范围的关联研究。当连接多项研究的数据时,研究参与者完成的特定研究和分级的选择加入同意可能不兼容或不明确。这阻止了用于回答新研究问题的协调的临床研究数据集。一些项目致力于开发同意本体,但目前还没有适合NHS收集同意指南的同意本体。莱斯特已经共同领导了全球基因组学与健康联盟在这一领域的努力,我将帮助扩展到一种基于本体的方法来代表NHS的同意和数据使用条件,以允许同意协调符合英国新数据保护法的要求。协调的数据集可以连接到标准化数据的公共来源,以弥合与翻译研究的差距。这方面的一个例子是在人类和老鼠的表型本体论之间映射的跨学科合作,以允许发现一组人类表型异常的老鼠疾病模型。如果一种疾病没有已知的遗传原因,进行跨物种表型比较的能力可以使这种疾病的潜在小鼠基因敲除模型被发现。这些跨物种映射已经应用到公共标准化数据库中,我将用真实世界与健康相关的大数据来调查它们的效用。
英文摘要
Connecting health-related research big data enables them to be compared, and increased sample/participant sizes to be discovered for analysis. Big data pose integration challenges with regards to their complexity. Phenotype is a data type that shows particular variety and variability across clinical and non-clinical (e.g. 'omics) health research big data. The term "phenotype" is used to define an aggregated set of medically and semantically distinct concepts such as a trait (e.g. blood glucose level), medical signs and symptoms (e.g. hyperglycemia), and disease (e.g. type 2 diabetes). To be able to compare values across big data we need to know if the values have the same meaning between datasets and the semantic rigour required to do this is provided by the use of "ontologies". There are several ontologies that describe overlapping phenotype domains but have been developed for different purposes, for example SNOMED CT is used by the NHS and the Human Phenotype Ontology (HPO) is used by research databases. If datasets are coded to different ontologies, or no ontology at all, then they can be linked to a common ontology via a process of harmonisation. This involves mapping terms from different ontologies, and text mining (TM) of free-text to associate the original values with ontology codes. Unfortunately, there are several barriers to this such as gaps in the ontologies and the publically available ontology mappings, the need to adapt current TM approaches which are optimised to perform well in clearly defined areas, and the need to scale current mapping and TM methods to work with big data.During this fellowship I will create enhanced capabilities for connecting clinical and non-clinical research big data by using ontologies to harmonise phenotype data. This will involve bridging the gaps in current ontologies, and adapting current state of the art TM and ontology mapping approaches so they are optimised for this context and can be applied to big data. The approaches I develop will be disease agnostic, however in the first instance they will be applied to disease areas of local interest. The Leicester Biomedical Research Centre focuses on cardiovascular, respiratory and lifestyle diseases and encompasses datasets from primary and secondary care, and clinical research studies which include participant questionnaires and biological sample data. I will connect clinical research data within and between disease areas and with local and publically available 'omics data, for example genome- and epigenome-wide association studies.The study-specific and tiered opt-in consent completed by study participants can be incompatible or ambiguous when connecting data across multiple studies. This blocks a harmonised clinical research dataset being used to answer a new research question. Some projects have worked on developing consent ontologies, but there is not currently a suitable consent ontology that fits with NHS guidance on collecting consent. Leicester already co-leads the Global Alliance for Genomics and Health efforts in this area, which I will help extend towards an ontology-based approach for representing NHS consents and data use conditions, to allow consent harmonisation in line with the requirements of new UK data protection laws.Harmonised datasets can be connected to public sources of standardised data to bridge the gap to translational research. An example of this are cross-disciplinary collaborations that have mapped between human and mouse phenotype ontologies, to allow the discovery of mouse disease models for a collection of human phenotypic abnormalities. Where a disease does not have a known genetic cause, the ability to perform a cross-species phenotype comparison allows potential mouse gene-knockout models for the disease to be discovered. These cross-species mappings have been applied to public standardised databases and I will investigate their utility with real-world health related big data.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1101/2023.06.23.546229
发表时间:
2023-06
期刊:
Journal of Proteome Research
影响因子:
4.4
作者:
[Meiqi Wang;Avish Vijayaraghavan;Tim Beck;J. Posma]
通讯作者:
Meiqi Wang;Avish Vijayaraghavan;Tim Beck;J. Posma
DOI:
10.1002/humu.24369
发表时间:
2022-06
期刊:
HUMAN MUTATION
影响因子:
3.9
作者:
[Rambla, Jordi, Baudis, Michael, Ariosa, Roberto, Beck, Tim, Fromont, Lauren A., Navarro, Arcadi, Paloots, Rahel, Rueda, Manuel, Saunders, Gary, Singh, Babita, Spalding, John D., Tornroos, Juha, Vasallo, Claudia, Veal, Colin D., Brookes, Anthony J.]
通讯作者:
Brookes, Anthony J.
DOI:
10.1016/j.xgen.2021.100029
发表时间:
2021-11-10
期刊:
CELL GENOMICS
影响因子:
--
作者:
[Rehm, Heidi L., Page, Angela J. H., Birney, Ewan]
通讯作者:
Birney, Ewan
DOI:
10.1101/2022.02.22.481457
发表时间:
2022-02
期刊:
Metabolites
影响因子:
4.1
作者:
[Cheng S. Yeung;Tim Beck;J. Posma]
通讯作者:
Cheng S. Yeung;Tim Beck;J. Posma
DOI:
10.1016/j.jlr.2023.100471
发表时间:
2023-12
期刊:
JOURNAL OF LIPID RESEARCH
影响因子:
6.5
作者:
[Price, Tara R., Emfinger, Christopher H., Schueler, Kathryn L., King, Sarah, Nicholson, Rebekah, Beck, Tim, Yandell, Brian S., Summers, Scott A., Holland, William L., Krauss, Ronald M., Keller, Mark P., Attie, Alan D.]
通讯作者:
Attie, Alan D.
共 7 条
FAIRClinical: FAIR-ification of Supplementary Data to Support Clinical Research
-
批准号:EP/Y036395/1
-
项目类别:Research Grant
-
资助金额:$12.9万
-
财政年份:2024
-
负责人:Tim Beck
-
依托单位:
海外基金