Trackerless Freehand Ultrasound with Sequence Modelling and Auxiliary Transformation Over Past and Future Frames

Trackerless Freehand Ultrasound with Sequence Modelling and Auxiliary Transformation Over Past and Future Frames
复制标题

DOI:
10.1109/isbi53787.2023.10230773
复制
发表时间:
2022-11
期刊:
2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI)
影响因子:
--
通讯作者:
Qi Li;Ziyi Shen;Qian Li;D. Barratt;T. Dowrick;M. Clarkson;Tom Kamiel Magda Vercauteren;Yipeng Hu
Qi Li;Ziyi Shen;Qian Li;D. Barratt;T. Dowrick;M. Clarkson;Tom Kamiel Magda Vercauteren;Yipeng Hu
中科院分区:
其他
文献类型:
--
作者:
Qi Li;Ziyi Shen;Qian Li;D. Barratt;T. Dowrick;M. Clarkson;Tom Kamiel Magda Vercauteren;Yipeng Hu

文献摘要

相似文献

在许多临床应用中,没有跟踪器的三维(3D)徒手超声(US)重建比其二维或跟踪的同行更有优势。在本文中,我们建议使用前馈和递归神经网络(rnn)来估计过去和未来2D图像中US帧之间的3D空间变换。利用临时可用的帧,进一步提出了一种多任务学习算法,利用它们之间大量的辅助转换预测任务。在一项志愿者研究中,对19名志愿者的38个前臂进行了228次扫描,获得了4万多帧美国帧,通过帧预测精度、体积重建重叠、累积跟踪误差和最终漂移来量化测试性能,并基于光学跟踪器的地面真实度。结果表明建模时空相关的输入帧以及输出转换的重要性,并且由于额外的过去和/或未来帧而进一步改进。表现最好的模型与预测中等间隔帧之间的转换有关,间隔小于10帧,每秒20帧(fps)。无论使用或不使用基于lstm的rnn,在距离预测转换超过一秒的地方添加帧几乎没有什么好处。有趣的是,使用所建议的方法,可能不再需要显式的序列内损失,以鼓励组合转换中的一致性或最小化累积错误。执行代码和志愿者数据将公开,以确保可重复性和进一步的研究。
Three-dimensional (3D) freehand ultrasound (US) reconstruction without a tracker can be advantageous over its two-dimensional or tracked counterparts in many clinical applications. In this paper, we propose to estimate 3D spatial transformation between US frames from both past and future 2D images, using feed-forward and recurrent neural networks (RNNs). With the temporally available frames, a further multi-task learning algorithm is proposed to utilise a large number of auxiliary transformation-predicting tasks between them. Using more than 40,000 US frames acquired from 228 scans on 38 forearms of 19 volunteers in a volunteer study, the hold-out test performance is quantified by frame prediction accuracy, volume reconstruction overlap, accumulated tracking error and final drift, based on ground-truth from an optical tracker. The results show the importance of modelling the temporal-spatially correlated input frames as well as output transformations, with further improvement owing to additional past and/or future frames. The best performing model was associated with predicting transformation between moderately-spaced frames, with an interval of less than ten frames at 20 frames per second (fps). Little benefit was observed by adding frames more than one second away from the predicted transformation, with or without LSTM-based RNNs. Interestingly, with the proposed approach, explicit within-sequence loss that encourages consistency in composing transformations or minimises accumulated error may no longer be required. The implementation code and volunteer data1 will be made publicly available ensuring reproducibility and further research.