Cleaning by clustering: methodology for addressing data quality issues in biomedical metadata.
Cleaning by clustering: methodology for addressing data quality issues in biomedical metadata.
复制标题
聚类清理:解决生物医学元数据中数据质量问题的方法
DOI:
10.1186/s12859-017-1832-4
复制
发表时间:
2017-09-18
影响因子:
3
通讯作者:
Dumontier M
中科院分区:
文献类型:
--
作者:
Hu W;Zaveri A;Qiu H;Dumontier M
The ability to efficiently search and filter datasets depends on access to high quality metadata. While most biomedical repositories require data submitters to provide a minimal set of metadata, some such as the Gene Expression Omnibus (GEO) allows users to specify additional metadata in the form of textual key-value pairs (e.g. sex: female). However, since there is no structured vocabulary to guide the submitter regarding the metadata terms to use, consequently, the 44,000,000+ key-value pairs in GEO suffer from numerous quality issues including redundancy, heterogeneity, inconsistency, and incompleteness. Such issues hinder the ability of scientists to hone in on datasets that meet their requirements and point to a need for accurate, structured and complete description of the data. In this study, we propose a clustering-based approach to address data quality issues in biomedical, specifically gene expression, metadata. First, we present three different kinds of similarity measures to compare metadata keys. Second, we design a scalable agglomerative clustering algorithm to cluster similar keys together. Our agglomerative cluster algorithm identified metadata keys that were similar, based on (i) name, (ii) core concept and (iii) value similarities, to each other and grouped them together. We evaluated our method using a manually created gold standard in which 359 keys were grouped into 27 clusters based on six types of characteristics: (i) age, (ii) cell line, (iii) disease, (iv) strain, (v) tissue and (vi) treatment. As a result, the algorithm generated 18 clusters containing 355 keys (four clusters with only one key were excluded). In the 18 clusters, there were keys that were identified correctly to be related to that cluster, but there were 13 keys which were not related to that cluster. We compared our approach with four other published methods. Our approach significantly outperformed them for most metadata keys and achieved the best average F-Score (0.63). Our algorithm identified keys that were similar to each other and grouped them together. Our intuition that underpins cleaning by clustering is that, dividing keys into different clusters resolves the scalability issues for data observation and cleaning, and keys in the same cluster with duplicates and errors can easily be found. Our algorithm can also be applied to other biomedical data types.
登录
查看更多内容
影响因子:
17.1
作者:
Sweeney TE;Shidham A;Wong HR;Khatri P
通讯作者:
Khatri P
影响因子:
5.8
作者:
Bodenhofer, Ulrich;Kothmeier, Andreas;Hochreiter, Sepp
通讯作者:
Hochreiter, Sepp
影响因子:
2.5
作者:
Hu, Wei;Qu, Yuzhong;Cheng, Gong
通讯作者:
Cheng, Gong
影响因子:
3
作者:
Morris JH;Apeltsin L;Newman AM;Baumbach J;Wittkop T;Su G;Bader GD;Ferrin TE
通讯作者:
Ferrin TE
影响因子:
14.9
作者:
Barrett T;Wilhite SE;Ledoux P;Evangelista C;Kim IF;Tomashevsky M;Marshall KA;Phillippy KH;Sherman PM;Holko M;Yefanov A;Lee H;Zhang N;Robertson CL;Serova N;Davis S;Soboleva A
通讯作者:
Soboleva A