Interrater agreement between American and Chinese sleep centers according to the 2014 AASM standard

Interrater agreement between American and Chinese sleep centers according to the 2014 AASM standard
复制标题

根据2014年AASM标准,美国和中国睡眠中心之间的评估者协议

DOI:
10.1007/s11325-019-01801-x
复制
发表时间:
2019-06-01
影响因子:
2.5
通讯作者:
Xu, Yan
Xu, Yan
中科院分区:
医学4区
文献类型:
--
作者:
Deng, Shujian;Zhang, Xin;Xu, Yan

文献摘要

被引文献

相似文献

目的使用2014年美国睡眠医学学会(AASM)手册确定睡眠阶段评分的实验室间可靠性。方法对40例不同样本的整夜多导睡眠图(PSG)进行分析。PSG被分割成37,642个30-s时期。中国的五名医生和美国的两名医生根据2014年AASM标准进行了评分。使用Cohen kappa(kappa)评价两个中心之间的评分不一致性。结果实验室间的信度有显著性差异(kappa=0.75 +/-0.01)。W期(kappa=0.89)和R期(kappa=0.87)的评分达到了最高一致性,而N1期(kappa=0.45)的评分反映了最低一致性。从相对不一致率来看,N2-N3(22.09%)、W-N1(19.68%)和N1-N2(18.75%)是最常见的不一致组合。中美医生在W-N1、N1-N2、N2-N3三种差异组合的评分上有一定的特点。分歧的原因有七个,即“阈上特性”(29.21%)、“语境影响”(18.06%)、“特征识别困难”(8.81%),“觉醒-清醒混淆”(7.57%),“推导不一致”(2.15%),“边缘特征”结论本研究证实了2014年AASM睡眠阶段评分的一致性,并探讨了标签模糊的潜在来源。因此,提出了改进措施,以帮助消除模糊的评分,提高评分的可靠性,在国际水平。
Objectives To determine inter-lab reliability in sleep stage scoring using the 2014 American Academy of Sleep Medicine (AASM) manual. To understand in-depth reasons for disagreement and provide suggestions for improvement.Methods This study consisted of 40 all-night polysomnographys (PSGs) from different samples. PSGs were segmented into 37,642 30-s epochs. Five doctors from China and two doctors from America scored the epochs following the 2014 AASM standard. Scoring disagreement between two centers was evaluated using Cohen's kappa (kappa). After visual inspection of PSGs of deviating scorings, potential disagreement reasons were analyzed.Results Inter-lab reliability yielded a substantial degree (kappa=0.75 +/- 0.01). Scoring for stage W (kappa=0.89) and R (kappa=0.87) achieved the highest agreement, while stage N1 (kappa=0.45) reflected the lowest. Considering the relative disagreement ratio, N2-N3 (22.09%), W-N1 (19.68%), and N1-N2 (18.75%) were the most frequent combinations of discrepancy. American and Chinese doctors showed certain characteristics in the scoring of discrepancy combination W-N1, N1-N2, and N2-N3. There are seven reasons for disagreement, namely "on-threshold characteristic" (29.21%), "context influence" (18.06%), "characteristic identification difficulty" (8.81%), "arousal-wake confusion" (7.57%), "derivation inconsistence" (2.15%), "on-borderline characteristic" (0.92%), and "misrecognition" (33.27%).Conclusions This study demonstrated the sleep stage scoring agreement of the 2014 AASM manual and explored potential sources of labeling ambiguity. Improvement measures were suggested accordingly to help remove ambiguity for scorers and improve scoring reliability at the international level.