Automatic identification of variables in epidemiological datasets using logic regression

Automatic identification of variables in epidemiological datasets using logic regression
复制标题

DOI:
10.1186/s12911-017-0429-1
复制
发表时间:
2017-04-13
影响因子:
3.5
通讯作者:
Orth, Andreas
Orth, Andreas
中科院分区:
医学3区
文献类型:
--
作者:
Lorenz, Matthias W.;Abdi, Negin Ashtiani;Orth, Andreas

文献摘要

被引文献

相似文献

背景资料:对于单个参与者数据(IPD)荟萃分析,多个数据集必须以一致的格式进行转换,例如使用统一的变量名称。当必须处理大量数据集时,这可能是一项耗时且容易出错的任务。变量的自动或半自动识别有助于减少工作量并提高数据质量。对于半自动化的高灵敏度在识别匹配的变量是特别重要的,因为它允许创建软件,其中的目标变量提出了一个选择的源变量,从中用户可以选择匹配的一个,只有低风险的错过了一个正确的源variable.Methods:对于每个变量中的一组目标变量,一些简单的规则手动创建。通过逻辑回归,使用流行病学和临床队列数据的大型数据库的随机子集(构造子集),为每个目标变量搜索这些规则的最佳布尔组合。在该数据库的第二个子集(验证子集)中,此最优组合rules进行了validated.Results:在构建样本中,平均分配了41个目标变量,阳性预测值(PPV)为34%,阴性预测值(NPV)为95%。在验证样品中,PPV为33%,而NPV保持在94%。在建设样本中,PPV是50%或更少的63%的所有变量,在验证样本中的71%的所有variables.Conclusions:我们证明了应用逻辑回归在一个复杂的数据管理任务,在大型流行病学IPD荟萃分析是可行的。然而,该算法的性能较差,这可能需要备份策略。
Background: For an individual participant data (IPD) meta-analysis, multiple datasets must be transformed in a consistent format, e.g. using uniform variable names. When large numbers of datasets have to be processed, this can be a time-consuming and error-prone task. Automated or semi-automated identification of variables can help to reduce the workload and improve the data quality. For semi-automation high sensitivity in the recognition of matching variables is particularly important, because it allows creating software which for a target variable presents a choice of source variables, from which a user can choose the matching one, with only low risk of having missed a correct source variable.Methods: For each variable in a set of target variables, a number of simple rules were manually created. With logic regression, an optimal Boolean combination of these rules was searched for every target variable, using a random subset of a large database of epidemiological and clinical cohort data (construction subset). In a second subset of this database (validation subset), this optimal combination rules were validated.Results: In the construction sample, 41 target variables were allocated on average with a positive predictive value (PPV) of 34%, and a negative predictive value (NPV) of 95%. In the validation sample, PPV was 33%, whereas NPV remained at 94%. In the construction sample, PPV was 50% or less in 63% of all variables, in the validation sample in 71% of all variables.Conclusions: We demonstrated that the application of logic regression in a complex data management task in large epidemiological IPD meta-analyses is feasible. However, the performance of the algorithm is poor, which may require backup strategies.