Cleaning by clustering: methodology for addressing data quality issues in biomedical metadata.

Cleaning by clustering: methodology for addressing data quality issues in biomedical metadata.
复制标题

聚类清理:解决生物医学元数据中数据质量问题的方法

DOI:
10.1186/s12859-017-1832-4
复制
发表时间:
2017-09-18
期刊:
影响因子:
3
通讯作者:
Dumontier M
Dumontier M
中科院分区:
生物学4区
文献类型:
--
作者:
Hu W;Zaveri A;Qiu H;Dumontier M

文献摘要

参考文献

被引文献

相似文献

有效搜索和过滤数据集的能力取决于对高质量元数据的访问。虽然大多数生物医学存储库要求数据提交者提供最小的元数据集,但有些数据库(如基因表达综合数据库(GEO))允许用户以文本键值对的形式指定额外的元数据(例如性别:女性)。然而,由于没有结构化的词汇表来指导子实体使用元数据术语,因此,GEO中的44,000,000多个键值对存在许多质量问题,包括冗余、异构、不一致和不完整。这些问题阻碍了科学家研究满足其要求的数据集的能力,并指出需要对数据进行准确,结构化和完整的描述。在这项研究中,我们提出了一种基于聚类的方法来解决生物医学,特别是基因表达,元数据的数据质量问题。首先,我们提出了三种不同的相似性度量来比较元数据键。其次,我们设计了一个可扩展的凝聚聚类算法聚类相似的关键在一起。我们的凝聚聚类算法基于(i)名称、(ii)核心概念和(iii)值相似性来识别彼此相似的元数据键,并将它们分组在一起。我们使用手动创建的金标准评估了我们的方法,其中359个关键词基于六种类型的特征分为27个簇:(i)年龄,(ii)细胞系,(iii)疾病,(iv)菌株,(v)组织和(vi)治疗。结果,该算法生成了包含355个关键字的18个聚类(排除了只有一个关键字的四个聚类)。在这18个聚类中,有一些键被正确识别为与该聚类相关,但有13个键与该聚类无关。我们将我们的方法与其他四种已发表的方法进行了比较。我们的方法在大多数元数据键上的表现明显优于它们,并获得了最好的平均F-Score(0.63)。我们的算法识别出彼此相似的密钥,并将它们分组在一起。我们的直觉是,通过聚类进行清洗的基础是,将键划分到不同的簇中解决了数据观察和清洗的可扩展性问题,并且可以很容易地找到相同簇中的重复和错误键。我们的算法也可以应用于其他生物医学数据类型。
The ability to efficiently search and filter datasets depends on access to high quality metadata. While most biomedical repositories require data submitters to provide a minimal set of metadata, some such as the Gene Expression Omnibus (GEO) allows users to specify additional metadata in the form of textual key-value pairs (e.g. sex: female). However, since there is no structured vocabulary to guide the submitter regarding the metadata terms to use, consequently, the 44,000,000+ key-value pairs in GEO suffer from numerous quality issues including redundancy, heterogeneity, inconsistency, and incompleteness. Such issues hinder the ability of scientists to hone in on datasets that meet their requirements and point to a need for accurate, structured and complete description of the data. In this study, we propose a clustering-based approach to address data quality issues in biomedical, specifically gene expression, metadata. First, we present three different kinds of similarity measures to compare metadata keys. Second, we design a scalable agglomerative clustering algorithm to cluster similar keys together. Our agglomerative cluster algorithm identified metadata keys that were similar, based on (i) name, (ii) core concept and (iii) value similarities, to each other and grouped them together. We evaluated our method using a manually created gold standard in which 359 keys were grouped into 27 clusters based on six types of characteristics: (i) age, (ii) cell line, (iii) disease, (iv) strain, (v) tissue and (vi) treatment. As a result, the algorithm generated 18 clusters containing 355 keys (four clusters with only one key were excluded). In the 18 clusters, there were keys that were identified correctly to be related to that cluster, but there were 13 keys which were not related to that cluster. We compared our approach with four other published methods. Our approach significantly outperformed them for most metadata keys and achieved the best average F-Score (0.63). Our algorithm identified keys that were similar to each other and grouped them together. Our intuition that underpins cleaning by clustering is that, dividing keys into different clusters resolves the scalability issues for data observation and cleaning, and keys in the same cluster with duplicates and errors can easily be found. Our algorithm can also be applied to other biomedical data types.
DOI: 10.1126/scitranslmed.aaa5993
发表时间: 2015-05-13
影响因子: 17.1
作者:
Sweeney TE;Shidham A;Wong HR;Khatri P
通讯作者: Khatri P
DOI: 10.1093/bioinformatics/btr406
发表时间: 2011-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Bodenhofer, Ulrich;Kothmeier, Andreas;Hochreiter, Sepp
通讯作者: Hochreiter, Sepp
匹配大型本体:分而治之的方法
DOI: 10.1016/j.datak.2008.06.003
发表时间: 2008-10-01
影响因子: 2.5
作者:
Hu, Wei;Qu, Yuzhong;Cheng, Gong
通讯作者: Cheng, Gong
clusterMaker:Cytoscape 的多算法聚类插件。
DOI: 10.1186/1471-2105-12-436
发表时间: 2011-11-09
期刊: BMC bioinformatics
影响因子: 3
作者:
Morris JH;Apeltsin L;Newman AM;Baumbach J;Wittkop T;Su G;Bader GD;Ferrin TE
通讯作者: Ferrin TE
DOI: 10.1093/nar/gks1193
发表时间: 2013-01
影响因子: 14.9
作者:
Barrett T;Wilhite SE;Ledoux P;Evangelista C;Kim IF;Tomashevsky M;Marshall KA;Phillippy KH;Sherman PM;Holko M;Yefanov A;Lee H;Zhang N;Robertson CL;Serova N;Davis S;Soboleva A
通讯作者: Soboleva A