Women's features and inter-/intra-rater agreement on mammographic density assessment in full-field digital mammograms (DDM-SPAIN)

Women's features and inter-/intra-rater agreement on mammographic density assessment in full-field digital mammograms (DDM-SPAIN)
复制标题

DOI:
10.1007/s10549-011-1833-3
复制
发表时间:
2012-02-01
影响因子:
3.8
通讯作者:
Salas, Dolores
Salas, Dolores
中科院分区:
医学2区
文献类型:
--
作者:
Perez-Gomez, Beatriz;Ruiz, Franciso;Salas, Dolores

文献摘要

被引文献

相似文献

乳腺摄影密度(MD)是乳腺癌的主要危险因素之一,其测量仍然依赖于主观评估。然而,MD测量在全数字乳腺X线摄影中的一致性尚未得到评价。我们研究了评估者之间和评估者内部对全数字化乳房X线照片中乳房密度估计的一致性,并测试了女性的任何特征是否会对她们产生影响。经过初步培训后,三名经验丰富的放射科医生使用博伊德量表对1,431名妇女的左乳头尾位乳房X光片进行MD评估,这些妇女是在三个西班牙筛查中心招募的。随机选择的50张图像的亚组被读取两次,以估计短期内的评价者一致性。此外,由一名评估者在2年前对1,428张图像进行阅读,用于估计长期评估者内一致性。计算成对加权kappas和95% bootstrap置信区间。定义了二分变量,以识别任何评估者与其他评估者或他/她自己的评估不一致的乳房X线照片。使用多变量混合逻辑模型,包括中心作为随机效应项,并考虑到重复测量时,需要的分歧和妇女的特征之间的关联进行了测试。评估者间和评估者内一致性的所有二次加权Kappa值均为优秀(高于0.80)。研究女性的特征,即体重指数、胸罩尺寸、绝经、未产、哺乳或当前激素治疗,均与评估者间或评估者内不一致的风险较高无关。然而,评分者在高密度MD类别中分类的图像中差异更大,并且在低密度乳房X线照片中,评分者内评估的不一致性也更低。全视野数字乳腺X线摄影中MD评估的可靠性与原始或数字化图像相当。受试者的MD相关特征与一致性之间缺乏相关性,这表明该来源的偏倚不太可能。
Measurement of mammographic density (MD), one of the leading risk factors for breast cancer, still relies on subjective assessment. However, the consistency of MD measurement in full-digital mammograms has yet to be evaluated. We studied inter- and intra-rater agreement with respect to estimation of breast density in full-digital mammograms, and tested whether any of the women's characteristics might have some influence on them. After an initial training period, three experienced radiologists estimated MD using Boyd scale in a left breast cranio-caudal mammogram of 1,431 women, recruited at three Spanish screening centres. A subgroup of 50 randomly selected images was read twice to estimate short-term intra-rater agreement. In addition, a reading of 1,428 of the images, performed 2 years before by one rater, was used to estimate long-term intra-rater agreement. Pair-wise weighted kappas with 95% bootstrap confidence intervals were calculated. Dichotomous variables were defined to identify mammograms in which any rater disagreed with other raters or with his/her own assessment, respectively. The association between disagreement and women's characteristics was tested using multivariate mixed logistic models, including centre as a random-effects term, and taking into account repeated measures when required. All quadratic-weighted kappa values for inter- and intra-rater agreement were excellent (higher than 0.80). None of the studied women's features, i.e. body mass index, brassiere size, menopause, nulliparity, lactation or current hormonal therapy, was associated with higher risk of inter- or intra-rater disagreement. However, raters differed significantly more in images that were classified in the higher-density MD categories, and disagreement in intra-rater assessment was also lower in low-density mammograms. The reliability of MD assessment in full-field digital mammograms is comparable to that for original or digitised images. The reassuring lack of association between subjects' MD-related characteristics and agreement suggests that bias from this source is unlikely.