Journal of Graphics ›› 2026, Vol. 47 ›› Issue (4): 726-735.DOI: 10.11996/JG.j.2095-302X.2026040726
• Image Processing and Computer Vision • Previous Articles Next Articles
LI Yuhua, JIANG Shan(
), YANG Zhiyong, WANG Yuze, ZHOU Zeyang
Received:2025-12-31
Accepted:2026-05-07
Online:2026-08-31
Published:2026-08-31
Contact:
JIANG Shan
Supported by:CLC Number:
LI Yuhua, JIANG Shan, YANG Zhiyong, WANG Yuze, ZHOU Zeyang. Physically-augmented and depth synergized freehand 3D ultrasound reconstruction[J]. Journal of Graphics, 2026, 47(4): 726-735.
Add to citation manager EndNote|Ris|BibTeX
URL: http://www.txxb.com.cn/EN/10.11996/JG.j.2095-302X.2026040726
| 编码器分支 | 输入特征(维度) | 层级 | 操作 | 输出通道 | 特征图尺寸 |
|---|---|---|---|---|---|
| 分支A:灰度流 | 原始灰度图 | 1 | Conv2D(1→32, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 32 | H/2×W/2 |
| 2 | Conv2D(32→64, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 64 | H/4×W/4 | ||
| 3 | Conv2D(64→128, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 128 | H/8×W/8 | ||
| 分支B:边缘流 | Canny距离变换 | 1 | Conv2D(1→32, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 32 | H/2×W/2 |
| 2 | Conv2D(32→64, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 64 | H/4×W/4 | ||
| 3 | Conv2D(64→128, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 128 | H/8×W/8 | ||
| 分支C:运动流 | 双向光流张量 | 1 | Conv2D(4→32, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 32 | H/2×W/2 |
| 2 | Conv2D(32→64, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 64 | H/4×W/4 | ||
| 3 | Conv2D(64→128, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 128 | H/8×W/8 |
Table 1 Configuration of the CNN encoder for physical-augmentation features
| 编码器分支 | 输入特征(维度) | 层级 | 操作 | 输出通道 | 特征图尺寸 |
|---|---|---|---|---|---|
| 分支A:灰度流 | 原始灰度图 | 1 | Conv2D(1→32, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 32 | H/2×W/2 |
| 2 | Conv2D(32→64, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 64 | H/4×W/4 | ||
| 3 | Conv2D(64→128, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 128 | H/8×W/8 | ||
| 分支B:边缘流 | Canny距离变换 | 1 | Conv2D(1→32, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 32 | H/2×W/2 |
| 2 | Conv2D(32→64, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 64 | H/4×W/4 | ||
| 3 | Conv2D(64→128, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 128 | H/8×W/8 | ||
| 分支C:运动流 | 双向光流张量 | 1 | Conv2D(4→32, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 32 | H/2×W/2 |
| 2 | Conv2D(32→64, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 64 | H/4×W/4 | ||
| 3 | Conv2D(64→128, k=3, p=1) + BN + ReLU + MaxPool2D(2) | 128 | H/8×W/8 |
| 模型编号 | CNN层数 | 物理增强通道 | 物理先验正则损失 | Speed loss | FDR/%↓ | MEA(°)↓ | |||
|---|---|---|---|---|---|---|---|---|---|
| 1 | 1 | √ | √ | √ | 0.68±0.21 | 32.90±10.18 | 56.76±21.08 | 41.62±14.62 | 7.52±3.86 |
| 2 | 3 | × | × | √ | 0.64±0.21 | 31.32±10.60 | 52.72±17.65 | 38.41±12.47 | 8.15±5.44 |
| 3 | 3 | √ | √ | × | 0.52±0.19 | 25.52±9.13 | 47.20±17.93 | 35.12±12.31 | 5.13±2.85 |
| 4 | 3 | √ | × | √ | 0.31±0.20 | 16.11±9.94 | 29.33±18.63 | 21.17±13.10 | 3.32±1.40 |
| 5 | 3 | √ | √ | √ | 0.27±0.06 | 13.13±10.51 | 23.75±10.61 | 16.97±4.38 | 3.01±0.96 |
Table 2 Ablation study quantitative performance metrics on the TUS-REC-Challenge dataset
| 模型编号 | CNN层数 | 物理增强通道 | 物理先验正则损失 | Speed loss | FDR/%↓ | MEA(°)↓ | |||
|---|---|---|---|---|---|---|---|---|---|
| 1 | 1 | √ | √ | √ | 0.68±0.21 | 32.90±10.18 | 56.76±21.08 | 41.62±14.62 | 7.52±3.86 |
| 2 | 3 | × | × | √ | 0.64±0.21 | 31.32±10.60 | 52.72±17.65 | 38.41±12.47 | 8.15±5.44 |
| 3 | 3 | √ | √ | × | 0.52±0.19 | 25.52±9.13 | 47.20±17.93 | 35.12±12.31 | 5.13±2.85 |
| 4 | 3 | √ | × | √ | 0.31±0.20 | 16.11±9.94 | 29.33±18.63 | 21.17±13.10 | 3.32±1.40 |
| 5 | 3 | √ | √ | √ | 0.27±0.06 | 13.13±10.51 | 23.75±10.61 | 16.97±4.38 | 3.01±0.96 |
Fig. 3 Training and validation curves of ablated models ((a) Convergence curves of the distance loss component; (b) Convergence curves of the composite total loss)
| 输入帧数/N | FDR/%↓ | MEA/°↓ | 单帧耗时/ms |
|---|---|---|---|
| 5 | 21.34±9.52 | 4.52±2.15 | 8 |
| 10 | 16.97±4.38 | 3.01±0.96 | 12 |
| 15 | 16.21±5.10 | 2.85±1.02 | 25 |
Table 3 Influence of input frame number N on system performance and efficiency
| 输入帧数/N | FDR/%↓ | MEA/°↓ | 单帧耗时/ms |
|---|---|---|---|
| 5 | 21.34±9.52 | 4.52±2.15 | 8 |
| 10 | 16.97±4.38 | 3.01±0.96 | 12 |
| 15 | 16.21±5.10 | 2.85±1.02 | 25 |
| 窗口尺寸 | 金字塔层数 | 迭代次数 | |
|---|---|---|---|
| 5×5 | 3 | 5 | 0.42±0.15 |
| 7×7 | 3 | 5 | 0.27±0.06 |
| 9×9 | 3 | 5 | 0.35±0.09 |
| 7×7 | 5 | 5 | 0.26±0.12 |
Table 4 Comparison under different hyperparameters
| 窗口尺寸 | 金字塔层数 | 迭代次数 | |
|---|---|---|---|
| 5×5 | 3 | 5 | 0.42±0.15 |
| 7×7 | 3 | 5 | 0.27±0.06 |
| 9×9 | 3 | 5 | 0.35±0.09 |
| 7×7 | 5 | 5 | 0.26±0.12 |
| 组合方案 | FDR/%↓ | |||
|---|---|---|---|---|
| A | 0.1 | 0.5 | 0.55 | 35.20 |
| B | 0.5 | 0.2 | 0.27 | 16.97 |
| C | 0.8 | 0.05 | 0.35 | 22.40 |
Table 5 Comparative experiments of different weight combination schemes
| 组合方案 | FDR/%↓ | |||
|---|---|---|---|---|
| A | 0.1 | 0.5 | 0.55 | 35.20 |
| B | 0.5 | 0.2 | 0.27 | 16.97 |
| C | 0.8 | 0.05 | 0.35 | 22.40 |
| 数据集 | 网络 | FDR/%↓ | MEA/°↓ | |||
|---|---|---|---|---|---|---|
| TUS-REC-Challenge | ConvLSTM | 0.69±0.23* | 34.56±12.31* | 24.38±8.20 | 44.42±18.37* | 7.07±4.13* |
| DC2-Net | 0.55±0.21* | 26.57±10.04* | 32.74±10.08* | 33.30±15.10* | 4.81±4.06** | |
| CNN-OF | 0.94±0.42* | 46.43±12.12* | 24.25±10.51 | 34.19±17.46* | 11.31±6.00* | |
| Efficientnet | 0.45±0.22 * | 19.28±13.25** | 26.55±13.29 | 29.33±9.41 * | 5.82±3.88* | |
| ResNet | 0.70±0.15* | 34.00±12.37* | 24.23±14.18 | 43.86±16.98* | 7.07±4.13* | |
| ResNet+LSTM | 0.60±0.20* | 29.36±12.09* | 23.27±4.63 | 36.34±16.34* | 7.23±3.99* | |
| PD2BNet | 0.27±0.06 | 13.13±10.51 | 23.75±10.61 | 16.97±4.38 | 3.01±0.96 | |
| Freehand_US_data | ConvLSTM | 0.73±0.25* | 36.28±12.95* | 25.07±8.46 | 46.19±19.03* | 7.34±4.28* |
| DC2-Net | 0.59±0.23* | 28.41±10.72* | 28.26±10.54 | 35.12±15.87* | 5.06±4.21 ** | |
| CNN-OF | 0.98±0.45* | 48.36±12.78* | 25.73±11.02 | 35.97±17.64* | 11.67±6.12* | |
| Efficientnet | 0.49±0.24* | 21.05±13.87 ** | 27.32±13.76 | 31.25±9.68* | 6.09±4.03* | |
| ResNet | 0.74±0.17* | 35.73±12.89* | 25.19±14.65 | 45.63±17.29* | 7.31±4.25* | |
| PLPPI+ | 0.50±0.16 | — | — | 28.21±13.24 | — | |
| ResNet+LSTM | 0.64±0.21* | 31.22±12.57* | 25.59±4.82 | 38.15±16.73* | 7.48±4.12* | |
| FiMA+ | 0.48±0.19 | — | — | 25.30±12.03 | — | |
| NING等 + | 0.44±0.17 | — | — | 22.53±9.88 | — | |
| PD2BNet | 0.30±0.11 | 14.06±11.03 | 24.59±11.25 | 18.24±4.72 | 3.27±1.05 |
Table 6 Comparison with state-of-the-art methods on Freehand_US_data and TUS-REC-Challenge datasets
| 数据集 | 网络 | FDR/%↓ | MEA/°↓ | |||
|---|---|---|---|---|---|---|
| TUS-REC-Challenge | ConvLSTM | 0.69±0.23* | 34.56±12.31* | 24.38±8.20 | 44.42±18.37* | 7.07±4.13* |
| DC2-Net | 0.55±0.21* | 26.57±10.04* | 32.74±10.08* | 33.30±15.10* | 4.81±4.06** | |
| CNN-OF | 0.94±0.42* | 46.43±12.12* | 24.25±10.51 | 34.19±17.46* | 11.31±6.00* | |
| Efficientnet | 0.45±0.22 * | 19.28±13.25** | 26.55±13.29 | 29.33±9.41 * | 5.82±3.88* | |
| ResNet | 0.70±0.15* | 34.00±12.37* | 24.23±14.18 | 43.86±16.98* | 7.07±4.13* | |
| ResNet+LSTM | 0.60±0.20* | 29.36±12.09* | 23.27±4.63 | 36.34±16.34* | 7.23±3.99* | |
| PD2BNet | 0.27±0.06 | 13.13±10.51 | 23.75±10.61 | 16.97±4.38 | 3.01±0.96 | |
| Freehand_US_data | ConvLSTM | 0.73±0.25* | 36.28±12.95* | 25.07±8.46 | 46.19±19.03* | 7.34±4.28* |
| DC2-Net | 0.59±0.23* | 28.41±10.72* | 28.26±10.54 | 35.12±15.87* | 5.06±4.21 ** | |
| CNN-OF | 0.98±0.45* | 48.36±12.78* | 25.73±11.02 | 35.97±17.64* | 11.67±6.12* | |
| Efficientnet | 0.49±0.24* | 21.05±13.87 ** | 27.32±13.76 | 31.25±9.68* | 6.09±4.03* | |
| ResNet | 0.74±0.17* | 35.73±12.89* | 25.19±14.65 | 45.63±17.29* | 7.31±4.25* | |
| PLPPI+ | 0.50±0.16 | — | — | 28.21±13.24 | — | |
| ResNet+LSTM | 0.64±0.21* | 31.22±12.57* | 25.59±4.82 | 38.15±16.73* | 7.48±4.12* | |
| FiMA+ | 0.48±0.19 | — | — | 25.30±12.03 | — | |
| NING等 + | 0.44±0.17 | — | — | 22.53±9.88 | — | |
| PD2BNet | 0.30±0.11 | 14.06±11.03 | 24.59±11.25 | 18.24±4.72 | 3.27±1.05 |
| 网络 | Params/M | FLOPs/G | Inference/s |
|---|---|---|---|
| 文献[16] | 6.52 | 3.63 | 16.63 |
| ResNet | 11.18 | 10.71 | 31.21 |
| PD2BNet | 5.26 | 12.39 | 18.50 |
Table 7 Comparison of Model Complexity and Efficiency on TUS-REC-Challenge Dataset (224×224, 1 500 frames)
| 网络 | Params/M | FLOPs/G | Inference/s |
|---|---|---|---|
| 文献[16] | 6.52 | 3.63 | 16.63 |
| ResNet | 11.18 | 10.71 | 31.21 |
| PD2BNet | 5.26 | 12.39 | 18.50 |
Fig. 5 3D trajectory comparisons of typical failure cases ((a) Left arm parallel scanning (LH_Par_C); (b) Right arm perpendicular scanning (RH_Ver_C))
| [1] |
ZHOU R, GUO F M, AZARPAZHOOH M R, et al. A voxel-based fully convolution network and continuous max-flow for carotid vessel-wall-volume segmentation from 3D ultrasound images[J]. IEEE Transactions on Medical Imaging, 2020, 39(9): 2844-2855.
DOI PMID |
| [2] |
YANG X, DOU H R, HUANG R B, et al. Agent with warm start and adaptive dynamic termination for plane localization in 3D ultrasound[J]. IEEE Transactions on Medical Imaging, 2021, 40(7): 1950-1961.
DOI URL |
| [3] |
LI Y H, JIANG S, YANG Z Y, et al. Deformable brain pMRI-iUS registration based on higher-order cliques Markov random field[J]. Biomedical Signal Processing and Control, 2026, 113: 109146.
DOI URL |
| [4] |
PREVOST R, SALEHI M, JAGODA S, et al. 3D freehand ultrasound without external tracking using deep learning[J]. Medical Image Analysis, 2018, 48: 187-202.
DOI PMID |
| [5] |
MOZAFFARI M H, LEE W S. Freehand 3-D ultrasound imaging: a systematic review[J]. Ultrasound in Medicine & Biology, 2017, 43(10): 2099-2124.
DOI URL |
| [6] |
PU G, JIANG S, YANG Z Y, et al. A novel ultrasound probe calibration method for multimodal image guidance of needle placement in cervical cancer brachytherapy[J]. Physica Medica, 2022, 100: 81-89.
DOI URL |
| [7] |
GAO H T, HUANG Q H, XU X M, et al. Wireless and sensorless 3D ultrasound imaging[J]. Neurocomputing, 2016, 195: 159-171.
DOI URL |
| [8] | MIURA K, ITO K, AOKI T, et al. Pose estimation of 2D ultrasound probe from ultrasound image sequences using CNN and RNN[C]// The 2nd International Workshop on Advances in Simplifying Medical Ultrasound. Cham: Springer, 2021: 96-105. |
| [9] |
MERCIER L, LANG T, LINDSETH F, et al. A review of calibration techniques for freehand 3-D ultrasound systems[J]. Ultrasound in Medicine & Biology, 2005, 31(2): 143-165.
DOI URL |
| [10] |
SUN R, LIU C B, WANG W S, et al. UltrasOM: a mamba-based network for 3D freehand ultrasound reconstruction using optical flow[J]. Computer Methods and Programs in Biomedicine, 2025, 268: 108843.
DOI URL |
| [11] | PREVOST R, SALEHI M, SPRUNG J, et al. Deep learning for sensorless 3D freehand ultrasound imaging[C]// The 20th International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer, 2017: 628-636. |
| [12] |
GUO H T, CHAO H Q, XU S, et al. Ultrasound volume reconstruction from freehand scans without tracking[J]. IEEE Transactions on Biomedical Engineering, 2023, 70(3): 970-979.
DOI URL |
| [13] | MOHAMED F, SIANG C V. A survey on 3D ultrasound reconstruction techniques[M]//ACEVES-FERNANDEZ M A. Artificial Intelligence-Applications in Medicine and Biology. London: IntechOpen, 2019: 73-92. |
| [14] | FARNEBÄCK G. Two-frame motion estimation based on polynomial expansion[C]// The 13th Scandinavian Conference on Image Analysis. Cham: Springer, 2003: 363-370. |
| [15] | CIPOLLA R, GAL Y, KENDALL A. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics[C]// 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 7482-7491. |
| [16] |
LI Q, SHEN Z Y, LI Q, et al. Long-term dependency for 3D reconstruction of freehand ultrasound without external tracker[J]. IEEE Transactions on Biomedical Engineering, 2024, 71(3): 1033-1042.
DOI URL |
| [17] | LI Q, SHEN Z Y, YANG Q Y, et al. Nonrigid reconstruction of freehand ultrasound without a tracker[C]// The 27th International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer, 2024: 689-699. |
| [18] | LI Q, SHEN Z Y, LI Q, et al. Trackerless freehand ultrasound with sequence modelling and auxiliary transformation over past and future frames[C]// The 20th International Symposium on Biomedical Imaging. New York: IEEE Press, 2023: 1-5. |
| [19] |
LUO M Y, YANG X, WANG H Z, et al. RecON: Online learning for sensorless freehand 3D ultrasound reconstruction[J]. Medical Image Analysis, 2023, 87: 102810.
DOI URL |
| [20] | NING G C, LIANG H Y, ZHOU L, et al. Spatial position estimation method for 3D ultrasound reconstruction based on hybrid transfomers[C]// The 19th International Symposium on Biomedical Imaging. New York: IEEE Press, 2022: 1-5. |
| [21] | YAN Z N, YANG X, LUO M Y, et al. Fine-grained context and multi-modal alignment for freehand 3D ultrasound reconstruction[C]// The 27th International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer, 2024: 340-349. |
| [22] |
DOU Y M, MU F Z, LI Y, et al. Sensorless end-to-end freehand 3-D ultrasound reconstruction with physics-guided deep learning[J]. IEEE Transactions on Ultrasonics, Ferroelectrics, and Frequency Control, 2024, 71(11): 1514-1525.
DOI URL |
| [1] | WANG Yutao, YANG Chao, KUANG Liqun, YANG Xiaowen, HAN Xie, JIAO Shichao. Hierarchical alignment for zero-shot sketch-based 3D shape retrieval [J]. Journal of Graphics, 2026, 47(4): 844-853. |
| [2] | LI Xiumei, ZHOU Zhengxin, SUN Junmei. A multi-task collaborative image forgery detection framework assisted by localization branch [J]. Journal of Graphics, 2026, 47(3): 524-533. |
| [3] | WU Wenhuan, WANG Wenshu, WANG Shuao. Monocular depth estimation method with hierarchical dual-stream attention [J]. Journal of Graphics, 2026, 47(3): 553-563. |
| [4] | LU Dehui, SONG Zhuo, HUANG Zhichao, TIAN Shiyu, LI Huimin, TIAN Mao, DENG Yichuan. Subjective visual perception prediction of green construction sites based on TrueSkill ranking and deep learning [J]. Journal of Graphics, 2026, 47(3): 641-652. |
| [5] | YAN Kang, ZENG Li, GU Xiaoqing. Cross-domain structured deep dictionary learning for image classification [J]. Journal of Graphics, 2026, 47(2): 341-350. |
| [6] | PANG Min, LI Zhentang, ZHANG Yuan, CUI Xiaokang, XIONG Fengguang. 3D model reconstruction based on retrieval and deformation techniques [J]. Journal of Graphics, 2026, 47(2): 368-379. |
| [7] | DONG Wenyi, YANG Weidong, TANG Binghui, WANG Qi, XIAO Hongyu. Review of deep learning based methods for detecting focal liver lesions [J]. Journal of Graphics, 2026, 47(1): 1-16. |
| [8] | ZHAI Yongjie, WANG Zixuan, ZHANG Zhenqi, ZHOU Xunqi, WANG Qianming. A vehicle damage classification model incorporating dual attention and weighted dynamic convolution [J]. Journal of Graphics, 2026, 47(1): 17-28. |
| [9] | PAN Yuxuan, JIN Rui, LIU Yu, ZHANG Lin. Generative model based unsupervised multi-view stereo network [J]. Journal of Graphics, 2026, 47(1): 29-38. |
| [10] | JIU Mingyuan, WU Guowei, SONG Xuguang, LI Shupan, XU Mingliang. Image classification method based on uncertainty-driven smart reinforcement active learning [J]. Journal of Graphics, 2026, 47(1): 47-56. |
| [11] | YANG Biao, WANG Xue, GUAN Zheng, LONG Ping. BSD-YOLO: a small target vehicle detection method based on dynamic sparse attention and adaptive detection head [J]. Journal of Graphics, 2026, 47(1): 99-110. |
| [12] | JU Chen, DING Jiaxin, WANG Zexing, LI Guangzhao, GUAN Zhenxiang, ZHANG Changyou. Graph neural network-based method for approximating finite element shape functions [J]. Journal of Graphics, 2025, 46(6): 1161-1171. |
| [13] | YI Bin, ZHANG Libin, LIU Danying, TANG Jun, FANG Junjun, LI Wenqi. Prediction model of laser drilling ventilation rate in cigarette manufacturing process based on AMTA-Net [J]. Journal of Graphics, 2025, 46(6): 1224-1232. |
| [14] | BO Wen, JU Chen, LIU Weiqing, ZHANG Yan, HU Jingjing, CHENG Jinghan, ZHANG Changyou. Degradation-driven temporal modeling method for equipment maintenance interval prediction [J]. Journal of Graphics, 2025, 46(6): 1233-1246. |
| [15] | ZHAO Zhenbing, Ouyang Wenbin, FENG Shuo, LI Haopeng, MA Jun. A thermal image detection method for insulators incorporating within-class sparse prior knowledge and improved YOLOv8 [J]. Journal of Graphics, 2025, 46(6): 1247-1256. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||