Probabilistic Matching Approach to Link Deidentified Data from a Trauma Registry and a Traumatic Brain Injury Model System Center.
Probabilistic Matching Approach to Link Deidentified Data from a Trauma Registry and a Traumatic Brain Injury Model System Center.
复制标题
用于链接来自创伤登记处和创伤性脑损伤模型系统中心的去识别数据的概率匹配方法。
DOI:
10.1097/phm.0000000000000513
复制
发表时间:
2017
影响因子:
3
通讯作者:
Wagner,AmyKathleen
中科院分区:
文献类型:
--
作者:
Kesinger,MatthewRyan;Kumar,RajGopalan;Ritter,AnneConnelly;Sperry,JasonLee;Wagner,AmyKathleen
ObjectiveThere is no civilian traumatic brain injury database that captures patients in all settings of the care continuum. The linkage of such databases would yield valuable insight into possible care interventions. Thus, the objective of this article is to describe the creation of an algorithm used to link the Traumatic Brain Injury Model System (TBIMS) to trauma data in state and national trauma databases.DesignThe TBIMS data from a single center was randomly divided into two sets. One subset was used to generate a probabilistic linking algorithm to link the TBIMS data to the center's trauma registry. The other subset was used to validate the algorithm. Medical record numbers were obtained and used as unique identifiers to measure the quality of the linkage. Novel methods were used to maximize the positive predictive value.ResultsThe algorithm generation subset had 121 patients. It had a sensitivity of 88% and a positive predictive value of 99%. The validation subset consisted of 120 patients and had a sensitivity of 83% and a positive predictive value of 99%.ConclusionsThe probabilistic linkage algorithm can accurately link TBIMS data across systems of trauma care. Future studies can use this database to answer meaningful research questions regarding the long-term impact of the acute trauma complex on health care utilization and recovery across the care continuum in traumatic brain injury populations.BACKGROUNDRecord linkage is a powerful tool in the field of public health. 1, 2 Through computational means, two large independent data sets can be combined to increase data sharing and provide opportunities to answer research questions not possible with a single data set alone. The two forms of record linkage are (1) deterministic and (2) probabilistic linkage. Crucially, the decision on the type of record linkage to use is based on the presence, or absence, of a unique identifier common between the two data sets. In instances where a unique identifier exists between two data sets, like first and last name or social security number, subjects with exact matches on the linking variables are defined as matches. Using this exact matching criterion of a unique identifier is known as deterministic record linkage. However, in instances where a common unique identifier is not available, it is possible that data sets may still be linked through probabilistic means. In this case, common data elements in both data sets can be compared to assess the likelihood that two patients are the same, given equal values on a number of variables.