Could automated machine-learned MRI grading aid epidemiological studies of lumbar spinal stenosis? Validation within the Wakayama spine study

Could automated machine-learned MRI grading aid epidemiological studies of lumbar spinal stenosis? Validation within the Wakayama spine study
复制标题

DOI:
10.1186/s12891-020-3164-1
复制
发表时间:
2020-03-12
影响因子:
2.3
通讯作者:
Fairbank, Jeremy
Fairbank, Jeremy
中科院分区:
医学3区
文献类型:
--
作者:
Ishimoto, Yuyu;Jamaludin, Amir;Fairbank, Jeremy

文献摘要

被引文献

相似文献

背景MRI扫描已经彻底改变了腰椎管狭窄症(LSS)的临床诊断。然而,目前还没有达成共识,如何最好地分类MRI结果,这阻碍了发展强大的纵向流行病学研究的条件。我们开发并测试了一个自动化系统,用于流行病学研究中使用的中央LSS腰椎MRI扫描分级。方法使用来自大规模人群队列研究(和歌山脊柱研究)的MRI扫描,全部由脊柱外科医生分级,我们训练了一个自动化系统,将中央LSS分为骨和软组织边缘的四个等级:无,轻度,中度,重度。随后,我们在测试集中根据观察者的独立读数测试了自动评分,以调查可靠性和一致性。结果971例受试者的4855个腰椎间节段获得完整的轴位视图。机器使用4365个轴向视图进行学习(训练集),并对其余490个轴向视图进行分级(测试集)。评分的符合率为65.7%(322/490),信度(林氏相关系数)为0.73。在2.2%的扫描(11/490)中,分类差异为2,仅0.2%(1/490)的差异为3。分为“重度”与“无/轻度/中度”2组。符合率为94.1%(461/490),Kappa值为0.75。结论本研究表明,自动化系统可以“学习”分级中央LSS与参考标准的优秀性能。因此,SpineNet提供了在大规模流行病学研究中对LSS进行分级的潜力,这些研究涉及具有高度一致性和客观性的大量MRI脊柱数据。
Background MRI scanning has revolutionized the clinical diagnosis of lumbar spinal stenosis (LSS). However, there is currently no consensus as to how best to classify MRI findings which has hampered the development of robust longitudinal epidemiological studies of the condition. We developed and tested an automated system for grading lumbar spine MRI scans for central LSS for use in epidemiological research. Methods Using MRI scans from the large population-based cohort study (the Wakayama Spine Study), all graded by a spinal surgeon, we trained an automated system to grade central LSS in four gradings of the bone and soft tissue margins: none, mild, moderate, severe. Subsequently, we tested the automated grading against the independent readings of our observer in a test set to investigate reliability and agreement. Results Complete axial views were available for 4855 lumbar intervertebral levels from 971 participants. The machine used 4365 axial views to learn (training set) and graded the remaining 490 axial views (testing set). The agreement rate for gradings was 65.7% (322/490) and the reliability (Lin's correlation coefficient) was 0.73. In 2.2% of scans (11/490) there was a difference in classification of 2 and in only 0.2% (1/490) was there a difference of 3. When classified into 2 groups as 'severe' vs 'no/mild/moderate'. The agreement rate was 94.1% (461/490) with a kappa of 0.75. Conclusions This study showed that an automated system can "learn" to grade central LSS with excellent performance against the reference standard. Thus SpineNet offers potential to grade LSS in large-scale epidemiological studies involving a high volume of MRI spine data with a high level of consistency and objectivity.