Comparison of Sequence Variants and the Application in Electronic Medical Records

Comparison of Sequence Variants and the Application in Electronic Medical Records
复制标题

序列变异体的比较及其在电子病历中的应用

DOI:
10.1007/978-3-031-12426-6_10
复制
发表时间:
2022
期刊:
Proceeding of the 33rd International Conference on Database and Expert Systems Applications (DEXA2022), Part 2
影响因子:
--
通讯作者:
Haruo Yokota
Haruo Yokota
中科院分区:
--
文献类型:
--
作者:
Yuqing Li;Hieu Hanh Le;Ryosuke Matsuo;Tomoyoshi Yamazak;Kenji Araki;Haruo Yokota

文献摘要

被引文献

相似文献

序列变体是一种表示部分有序元素的数据结构,它可以被视为具有分支的序列。序列变量的比较是实际应用中的一个重要问题。为了进行比较,需要检测序列变异的常见部分和不常见部分,但目前还没有合适的方法。在本研究中,我们通过提供适当的定义和算法,开发了一种比较序列变体的方法。在最长公共子序列的基础上定义了最长公共子序列变体,并提出了一个合并的序列变体进行比较。我们还提出了计算这些序列变体的算法。作为一个例子,我们将这些方法应用于来自电子病历(emr)的真实数据,以确定不同医院治疗模式的多样性,然后将结果展示给医务工作者,以帮助他们认识到差异并改进他们的医疗行动。通过序列模式挖掘,从emr的真实治疗顺序数据库中提取频繁的治疗模式,计算治疗模式之间的差异,并将其可视化为最长公共子序列变体和合并序列变体。
A sequence variant is a data structure that represents elements with partial order, and it can be regarded as a sequence with branches. The comparison of sequence variants is an important problem in practical applications. To conduct comparisons, it is necessary to detect the common and uncommon parts of sequence variants, but a suitable method is not available. In this study, we developed a method for comparing sequence variants by providing appropriate definitions and algorithms. The longest common subsequence variant is defined based on the longest common subsequence and a merged sequence variant is proposed for comparison. We also propose algorithms for calculating these sequence variants. As an example, we applied the methods to real data from electronic medical records (EMRs) to determine the diversity of the treatment patterns in different hospitals, before presenting the results to medical workers to help them recognize the differences and improve their medical actions. Frequent treatment patterns were extracted from a real treatment order database of EMRs by using sequential pattern mining, and the differences in the treatment patterns were calculated and visualized as longest common subsequence variants and merged sequence variants.