欢迎访问《图学学报》

图学学报 ›› 2026, Vol. 47 ›› Issue (4): 844-853.DOI: 10.11996/JG.j.2095-302X.2026040844

• 数字化设计与制造 • 上一篇    下一篇

基于层次化对齐的三维模型零样本草图检索

王雨涛1,2,3, 杨超1, 况立群1,2,3, 杨晓文1,2,3, 韩燮1,2,3, 焦世超1,2,3()   

  1. 1 中北大学计算机科学与技术学院山西 太原 030051
    2 机器视觉与虚拟现实山西省重点实验室山西 太原 030051
    3 山西省视觉信息处理及智能机器人工程研究中心山西 太原 030051
  • 收稿日期:2026-03-14 接受日期:2026-05-22 出版日期:2026-08-31 发布日期:2026-08-31
  • 通讯作者:焦世超,E-mail:20230006@nuc.edu.cn
  • 基金资助:
    国家自然科学基金(62272426);山西省重点研发计划项目(202402020101001);山西省自然科学基金(202303021212189)

Hierarchical alignment for zero-shot sketch-based 3D shape retrieval

WANG Yutao1,2,3, YANG Chao1, KUANG Liqun1,2,3, YANG Xiaowen1,2,3, HAN Xie1,2,3, JIAO Shichao1,2,3()   

  1. 1 School of Computer Science and Technology, North University of China, Taiyuan Shanxi 030051, China
    2 Shanxi Key Laboratory of Machine Vision and Virtual Reality, Taiyuan Shanxi 030051, China
    3 Shanxi Province’s Vision Information Processing and Intelligent Robot Engineering Research Center, Taiyuan Shanxi 030051, China
  • Received:2026-03-14 Accepted:2026-05-22 Published:2026-08-31 Online:2026-08-31
  • Contact: JIAO Shichao,E-mail:20230006@nuc.edu.cn
  • Supported by:
    National Natural Science Foundation of China(62272426);Key Research and Development Program of Shanxi Province(202402020101001);Natural Science Foundation of Shanxi(202303021212189)

摘要:

基于草图的三维模型检索,由于其直观的人机交互方式,已成为三维模型检索领域中一个重要研究方向。然而,在实际应用中三维模型类别数量庞大且持续增长,训练数据难以覆盖所有潜在类别,使得传统基于类别标签的监督学习检索方法难以适应开放环境下的新类别检索需求。因此,基于草图的三维模型零样本检索成为提升模型开放场景适应能力的重要研究内容。该任务不仅需要解决草图和三维模型之间的模态差异,还需要在缺乏未知类别样本的条件下,实现从已见类别向未见类别的有效知识迁移。为解决上述问题,设计了层次化对齐方法,通过输入层与特征层的渐进式对齐策略缓解跨模态差异的同时实现知识迁移。在输入层,通过对抗学习方法生成伪视图,将草图转化为更接近三维模型多视图的伪视图表示,从源头缓解模态差异;在特征层,先建立共享的语义嵌入空间,结合多粒度度量学习方法,引入语义一致性约束、语义增强机制和类级代理约束,提升特征表达能力。在SHREC 2013和SHREC 2014数据集上的实验结果表明层次化对齐框架在多项性能指标上取得了具有竞争力的结果。消融实验进一步证明伪视图生成机制和多粒度度量学习策略的有效性,说明输入层的模态对齐以及特征层的语义对齐能够协同优化跨模态特征表示,提升零样本检索性能。

关键词: 基于草图的三维模型检索, 零样本学习, 深度学习, 对抗生成, 跨模态对齐

Abstract:

Sketch-based 3D shape retrieval has gradually become an important research direction in the field of 3D shape retrieval, mainly because it provides an intuitive mode of human-computer interaction for users. However, in real application scenarios, the number of 3D shape categories is very large and continues to grow, which makes it hard for training data to cover all potential categories, and thus traditional supervised retrieval methods based on category labels often struggle to meet the demand for new-category retrieval in open environments. This limitation motivates the study of zero-shot sketch-based 3D shape retrieval, which aims to improve the adaptability of retrieval systems to open scenarios. In this task, there are two main difficulties: addressing the modality gap between sketches and 3D shapes, and achieving effective knowledge transfer from seen categories to unseen categories in the absence of samples from unseen classes. To address these problems, a hierarchical alignment method was designed, which performs progressive alignment at both the input level and the feature level to jointly reduce cross-modal differences and support knowledge transfer. At the input level, an adversarial learning-based pseudo-view generation method was introduced to convert sketches into pseudo multi-view representations that are closer to 3D shape distributions, thereby reducing modality discrepancy from the source. At the feature level, a shared semantic embedding space was constructed, and a multi-granularity metric learning strategy was further introduced, in which semantic consistency constraints, semantic enhancement mechanisms, and class-level proxy constraints were jointly used to improve feature representation. Experimental results on the SHREC 2013 and SHREC 2014 datasets showed that the proposed hierarchical alignment framework achieved competitive performance across multiple evaluation metrics, while ablation studies further verified the effectiveness of each component, indicating that input-level alignment and feature-level semantic alignment can work together to enhance cross-modal representation learning and improve zero-shot retrieval performance.

Key words: sketch-based 3D shape retrieval, zero-shot learning, deep learning, adversarial generation, cross-modal alignment

中图分类号: