Journal of Graphics ›› 2026, Vol. 47 ›› Issue (4): 746-756.DOI: 10.11996/JG.j.2095-302X.2026040746
• Computer Graphics and Virtual Reality • Previous Articles Next Articles
LIU Qu1,3, CHEN Bin2,3(
), HUANG Yuanzheng1,3
Received:2026-04-15
Accepted:2026-05-04
Online:2026-08-31
Published:2026-08-31
Contact:
CHEN Bin
Supported by:CLC Number:
LIU Qu, CHEN Bin, HUANG Yuanzheng. QC-ORF: constructing query-conditioned object response fields in 3D Gaussians via weak prompts[J]. Journal of Graphics, 2026, 47(4): 746-756.
Add to citation manager EndNote|Ris|BibTeX
URL: http://www.txxb.com.cn/EN/10.11996/JG.j.2095-302X.2026040746
| 方法 | AUROC↑ | AUPRC↑ | gap↑ | ||
|---|---|---|---|---|---|
| 二维教师特征检索 | 0.818 5 | 0.520 7 | 0.500 8 | 0.498 4 | 0.002 4 |
| 硬阈值检索 | 0.957 0 | 0.909 9 | 0.898 3 | 0.453 1 | 0.445 3 |
| 单视角查询 | 0.992 5 | 0.947 9 | 0.926 3 | 0.331 7 | 0.594 6 |
| 多视角均值聚合 | 0.993 4 | 0.940 2 | 0.965 5 | 0.283 5 | 0.672 0 |
| 多视角加权聚合 | 0.991 3 | 0.948 4 | 0.987 1 | 0.252 3 | 0.734 8 |
Table 1 Quantitative comparison of QC-ORF
| 方法 | AUROC↑ | AUPRC↑ | gap↑ | ||
|---|---|---|---|---|---|
| 二维教师特征检索 | 0.818 5 | 0.520 7 | 0.500 8 | 0.498 4 | 0.002 4 |
| 硬阈值检索 | 0.957 0 | 0.909 9 | 0.898 3 | 0.453 1 | 0.445 3 |
| 单视角查询 | 0.992 5 | 0.947 9 | 0.926 3 | 0.331 7 | 0.594 6 |
| 多视角均值聚合 | 0.993 4 | 0.940 2 | 0.965 5 | 0.283 5 | 0.672 0 |
| 多视角加权聚合 | 0.991 3 | 0.948 4 | 0.987 1 | 0.252 3 | 0.734 8 |
| 方法 | Bear | Teatime | Ramen | ALL | |||
|---|---|---|---|---|---|---|---|
| mIoU | BmIoU | mIoU | BmIoU | mIoU | BmIoU | mIoU | |
| LERF[ | 57.2 | 52.6 | 44.1 | 40.6 | 27.4 | 14.9 | 45.6 |
| LangSplat[ | 74.4 | 71.8 | 64.8 | 62.1 | 50.9 | 47.1 | 64.5 |
| Gaussian Grouping[ | 82.6 | 80.4 | 70.3 | 65.7 | 74.5 | 70.7 | 68.4 |
| SAGA[ | 80.9 | 75.5 | 69.2 | 67.9 | 64.7 | 62.0 | 66.6 |
| Trace3D[ | 77.5 | 74.2 | 63.6 | 60.3 | 68.1 | 64.9 | 64.9 |
| Unified-Lift[ | 84.2 | 81.4 | 72.0 | 68.2 | 72.8 | 69.5 | 71.5 |
| QC-ORF | 86.4 | 83.7 | 70.6 | 67.4 | 73.1 | 71.2 | 72.4 |
Table 2 Comparison results on the LERF dataset /%
| 方法 | Bear | Teatime | Ramen | ALL | |||
|---|---|---|---|---|---|---|---|
| mIoU | BmIoU | mIoU | BmIoU | mIoU | BmIoU | mIoU | |
| LERF[ | 57.2 | 52.6 | 44.1 | 40.6 | 27.4 | 14.9 | 45.6 |
| LangSplat[ | 74.4 | 71.8 | 64.8 | 62.1 | 50.9 | 47.1 | 64.5 |
| Gaussian Grouping[ | 82.6 | 80.4 | 70.3 | 65.7 | 74.5 | 70.7 | 68.4 |
| SAGA[ | 80.9 | 75.5 | 69.2 | 67.9 | 64.7 | 62.0 | 66.6 |
| Trace3D[ | 77.5 | 74.2 | 63.6 | 60.3 | 68.1 | 64.9 | 64.9 |
| Unified-Lift[ | 84.2 | 81.4 | 72.0 | 68.2 | 72.8 | 69.5 | 71.5 |
| QC-ORF | 86.4 | 83.7 | 70.6 | 67.4 | 73.1 | 71.2 | 72.4 |
| 模块 | AUROC | AUPRC | gap |
|---|---|---|---|
| 全部模块 | 0.991 3 | 0.948 4 | 0.734 8 |
| w/o cosine loss | 0.536 5 | 0.357 9 | 0.214 2 |
| w/o inst loss | 0.921 4 | 0.852 8 | 0.654 2 |
| w/o 前景偏置 | 0.996 2 | 0.842 1 | 0.425 7 |
| w/o 负样本惩罚 | 0.788 4 | 0.601 0 | 0.266 3 |
| w/o 深度一致性门控 | 0.907 6 | 0.772 3 | 0.452 3 |
Table 3 Ablation studies of each module
| 模块 | AUROC | AUPRC | gap |
|---|---|---|---|
| 全部模块 | 0.991 3 | 0.948 4 | 0.734 8 |
| w/o cosine loss | 0.536 5 | 0.357 9 | 0.214 2 |
| w/o inst loss | 0.921 4 | 0.852 8 | 0.654 2 |
| w/o 前景偏置 | 0.996 2 | 0.842 1 | 0.425 7 |
| w/o 负样本惩罚 | 0.788 4 | 0.601 0 | 0.266 3 |
| w/o 深度一致性门控 | 0.907 6 | 0.772 3 | 0.452 3 |
| [1] | KERBL B, KOPANAS G, LEIMKÜHLER T, et al. 3D Gaussian splatting for real-time radiance field rendering[J]. ACM Transactions on Graphics, 2023, 42(4): 139. |
| [2] | MILDENHALL B, SRINIVASAN P P, TANCIK M, et al. NeRF: representing scenes as neural radiance fields for view synthesis[J]. Communications of the ACM, 2022, 65(1): 99-106. |
| [3] | YARIV L, GU J T, KASTEN Y, et al. Volume rendering of neural implicit surfaces[C]// The 35th International Conference on Neural Information Processing Systems. New York: ACM, 2021: 367. |
| [4] | BARRON J T, MILDENHALL B, TANCIK M, et al. Mip-NeRF: a multiscale representation for anti-aliasing neural radiance fields[C]// 2021 IEEE/CVF International Conference on Computer Vision. New York: IEEE Press, 2021: 5835-5844. |
| [5] | MÜLLER T, EVANS A, SCHIED C, et al. Instant neural graphics primitives with a multiresolution hash encoding[J]. ACM Transactions on Graphics, 2022, 41(4): 102. |
| [6] | QIN M H, LI W H, ZHOU J W, et al. LangSplat: 3D language Gaussian splatting[C]// 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2024: 20051-20060. |
| [7] | ZHOU S J, CHANG H R, JIANG S C, et al. Feature 3DGS: supercharging 3D Gaussian splatting to enable distilled feature fields[C]// 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2024: 21676-21685. |
| [8] |
ZUO X X, SAMANGOUEI P, ZHOU Y W, et al. FMGS: foundation model embedded 3D Gaussian splatting for holistic 3D scene understanding[J]. International Journal of Computer Vision, 2025, 133(2): 611-627.
DOI |
| [9] | GUO J, MA X J, FAN Y, et al. Semantic Gaussians:open-vocabulary scene understanding with 3D Gaussian splatting[EB/OL]. 2026. [2026-03-18]. https://doi.org/10.1109/TCSVT.2026.3675320. |
| [10] | YE M Q, DANELLJAN M, YU F, et al. Gaussian grouping: segment and edit anything in 3D scenes[C]// The 18th European Conference on Computer Vision. Cham: Springer, 2025: 162-179. |
| [11] | CEN J Z, FANG J M, YANG C, et al. Segment any 3D Gaussians[EB/OL]. [2026-02-15]. https://jumpat.github.io/SAGA/. |
| [12] | LAN K, LI H R, SHI H L, et al. 2D-guided 3D Gaussian segmentation[C]// 2024 Asian Conference on Communication and Networks. New York: IEEE Press, 2024: 1-5. |
| [13] | CHEN Y W, CHEN Z L, ZHANG C, et al. GaussianEditor: swift and controllable 3D editing with Gaussian splatting[C]// 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2024: 21476-21485. |
| [14] | GUÉDON A, LEPETIT V. SuGaR: surface-aligned Gaussian splatting for efficient 3D mesh reconstruction and high-quality mesh rendering[C]// 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2024: 5354-5363. |
| [15] | XIE T Y, ZONG Z S, QIU Y X, et al. PhysGaussian: physics-integrated 3D Gaussians for generative dynamics[C]// 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2024: 4389-4398. |
| [16] | KERR J, KIM C M, GOLDBERG K, et al. LERF: language embedded radiance fields[C]// 2023 IEEE/CVF International Conference on Computer Vision. New York: IEEE Press, 2023: 19672-19682. |
| [17] | LIU K H, ZHAN F N, ZHANG J H, et al. Weakly supervised 3D open-vocabulary segmentation[C]// The 37th International Conference on Neural Information Processing Systems. New York: ACM, 2023: 2325. |
| [18] | ZHANG H, LI F, AHUJA N. Open-NeRF: towards open vocabulary NeRF decomposition[C]// 2024 IEEE/CVF Winter Conference on Applications of Computer Vision. New York: IEEE Press, 2024: 3444-3453. |
| [19] | MIRZAEI A, AUMENTADO-ARMSTRONG T, DERPANIS K G, et al. SPIn-NeRF: multiview segmentation and perceptual inpainting with neural radiance fields[C]// 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2023: 20669-20679. |
| [20] | CEN J Z, ZHOU Z W, FANG J M, et al. Segment anything in 3D with NeRFs[C]// The 37th International Conference on Neural Information Processing Systems. New York: ACM, 2023: 1130. |
| [21] | RADFORD A, KIM J W, HALLACY C, et al. Learning transferable visual models from natural language supervision[EB/OL]. [2026-02-15]. https://dblp.uni-trier.de/db/conf/icml/icml2021.html#. |
| [22] | RAVI N, GABEUR V, HU Y T, et al. SAM 2:segment anything in images and videos[EB/OL]. [2025-12-28]. https://dblp.uni-trier.de/db/conf/iclr/iclr2025.html#conf/iclr/RaviGHHR0KRRGMP25. |
| [23] | SIMÉONI O, VO H V, SEITZER M, et al. DINOv3[EB/OL]. [2025-12-13]. https://arxiv.org/pdf/2508.10104. |
| [24] |
刘高屹, 胡瑞珍, 刘利刚. 基于2D特征蒸馏的3D高斯泼溅语义分割与编辑[J]. 图学学报, 2025, 46(2): 312-321.
DOI |
|
LIU G Y, HU R Z, LIU L G. 3D Gaussian splatting semantic segmentation and editing based on 2D feature distillation[J]. Journal of Graphics, 2025, 46(2): 312-321 (in Chinese).
DOI |
|
| [25] | CHEN K J, DAI B Q, QIN M H, et al. SLGaussian: fast language Gaussian splatting in sparse views[C]// The 33rd ACM International Conference on Multimedia. New York: ACM, 2025: 3047-3056. |
| [26] | GALERNE B, WANG J L, RAAD L, et al. SGSST: scaling Gaussian splatting style transfer[C]// 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2025: 26535-26544. |
| [27] | JATAVALLABHULA K M, KUWAJERWALA A, GU Q, et al. ConceptFusion:Open-set multimodal 3D mapping[EB/OL]. [2025-12-23]. https://dblp.uni-trier.de/db/conf/rss/rss2023.html#conf/rss/JatavallabhulaK23. |
| [28] | WU Y M, MENG J R, LI H J, et al. OpenGaussian: towards point-level 3D Gaussian-based open vocabulary understanding[C]// The 38th International Conference on Neural Information Processing Systems. New York: ACM, 2024: 604. |
| [29] | KOBAYASHI S, MATSUMOTO E, SITZMANN V. Decomposing nerf for editing via feature field distillation[C]// The 36th International Conference on Neural Information Processing Systems. New York: ACM, 2022: 1694. |
| [30] | Kim C M, Wu M X, Kerr J, et al. GARField: group anything with radiance fields[C]// 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2024: 21530-21539. |
| [31] | YING H Y, YIN Y X, ZHANG J Z, et al. OmniSeg3D: omniversal 3D segmentation via hierarchical contrastive learning[C]// 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2024: 20612-20622. |
| [32] | LEE H, YUN Y, BAE J, et al. Rethinking open-vocabulary segmentation of radiance fields in 3D space[C]// The 39th AAAI Conference on Artificial Intelligence. Washington, DC: AAAI Press, 2025: 4491-4498. |
| [33] | LIN H T, CHEN S L, LIEW J, et al. Depth anything 3:recovering the visual space from any views[EB/OL]. [2025-12-13]. https://arxiv.org/pdf/2511.10647. |
| [34] |
WANG Z, BOVIK A C, SHEIKH H R, et al. Image quality assessment: from error visibility to structural similarity[J]. IEEE Transactions on Image Processing, 2004, 13(4): 600-612.
DOI PMID |
| [35] | BARRON J T, MILDENHALL B, VERBIN D, et al. Mip-NeRF 360: unbounded anti-aliased neural radiance fields[C]// 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2022: 5460-5469. |
| [36] | REN T H, LIU S L, ZENG A L, et al. Grounded SAM: assembling open-world models for diverse visual tasks[EB/OL]. [2025-12-15]. https://arxiv.org/pdf/2401.14159. |
| [37] |
FAWCETT T. An introduction to ROC analysis[J]. Pattern Recognition Letters, 2006, 27(8): 861-874.
DOI URL |
| [38] | DAVIS J, GOADRICH M H. The relationship between precision-recall and ROC curves[EB/OL]. [2025-12-15]. https://dblp.uni-trier.de/db/conf/icml/icml2006.html#conf/icml/DavisG06. |
| [39] | SHEN H Y, NI J F, CHEN Y X, et al. Trace3D: consistent segmentation lifting via Gaussian instance tracing[C]// 2025 IEEE/CVF International Conference on Computer Vision. New York: IEEE Press, 2025: 6656-6666. |
| [40] | ZHU R S, QIU S, LIU Z Z, et al. Rethinking end-to-end 2D to 3D scene segmentation in Gaussian splatting[C]// 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2025: 3656-3665. |
| [1] | WANG Ziwei, WANG Lutao, LI Antong, SHEN Yan. Few-shot 3D Gaussian splatting based on monocular depth ambiguity-aware estimation [J]. Journal of Graphics, 2026, 47(4): 704-713. |
| [2] | XU Hang, XIE Xueguang, XIA Qing, GAO Yang, YU Peng, HU Jiahao. Gaussian dynamic reconstruction based on semantic perception and hybrid material point method [J]. Journal of Graphics, 2026, 47(4): 714-725. |
| [3] | LI Jitong, HE Jinxu, XUE Suling, ZHANG Jun, LOU Lu. Efficient 3D Gaussian splatting based on VGGT and saliency-guided voxelization [J]. Journal of Graphics, 2026, 47(3): 500-510. |
| [4] | LIAO Jiankang, ZHANG Yanci. Frequency intensity Gaussian splatting for over-reconstruction issue [J]. Journal of Graphics, 2026, 47(3): 564-575. |
| [5] | HUANG Jing, SHI Ruihao, SONG Wenming, GUO Hepan, WEI Huang, WEI Xiaosong, YAO Jian. A review of autonomous driving image synthesis methods: from simulators to new paradigms [J]. Journal of Graphics, 2025, 46(5): 931-949. |
| [6] | LIU Gaoyi, HU Ruizhen, LIU Ligang. 3D Gaussian splatting semantic segmentation and editing based on 2D feature distillation [J]. Journal of Graphics, 2025, 46(2): 312-321. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||